Videogame localization · audio QC

Automated Audio QC for Videogame Dubbing

An app that listens for you: names, durations, loudness and — line by line — whether each take actually says its script line. 1,000 takes cross-checked in 295 ms.

ASSETS CONTROL QC screen: take table with status chips, overlaid waveforms, script line versus what the take actually says at 100% match, and loudness metrics

Real product — the QC screen on the built-in demo project. Takes and script lines are sample material.

Videogame localization · audio QC

ASSETS CONTROL — Videogame Audio QC

1,000 takes cross-checked in 295 ms

Problem
A videogame dub arrives from the recording studio with hundreds of takes. Someone has to check every filename against the original, every duration against the cue sheet's per-line tolerance, every level against a spec that arrives as prose ("around −12 to −18 dB") — and listen to each take to confirm it actually says its line. Then rebuild the client's delivery folder tree by hand.
What we built
An app that listens for you. It transcribes every take on-device and compares it to the script line in its own language — catching takes that don't say their line, cross-named takes, wrong-language deliveries and half-read lines. It checks names, durations against the cue sheet's own tolerance column, and loudness in both dB and EBU R128. Then it builds the outgoing delivery with the original's folder structure, fills the AsRec column into the client's own Excel, and hands the operator a report that starts with the reds instead of confirming what's already fine. Everything runs locally: no API, no account, no internet — not one byte of client material leaves the machine.
What changed
The cross-check of a thousand takes now takes under a second, and loudness measures at ~115× real time. The operator's day inverted: instead of listening through everything to find the problems, the machine listens through everything and the operator starts from the problem list. And the honest before/after: checking to this level of detail simply wasn't done before — a 9,000-take delivery would have taken at least two people three weeks, and content-level listening still wouldn't have been attempted at this precision, by the owner's account.

Zero recurring cost — open models, on-device. Built for Caja de Ruidos' localization operation.

What the system does

  • Cross-checks the original delivery against the localized one: missing, extra, wrongly named or wrong-language files
  • Transcribes every take on-device and compares it word by word against its script line — in the line's own language
  • Catches crossed takes (“this is another file's line”), wrong-language deliveries and half-read lines
  • Reads per-line tolerances straight from the cue sheet — percentages, milliseconds, frames, LIP / ON / OFF / WILD
  • Loudness both ways studios ask for it: average/peak dB or EBU R128 LUFS with true peak, plus presets for common delivery specs
  • Technical checks: sample-rate and channel mismatches, clipped samples, head and tail silence, mute files
  • A/B listening from the keyboard with overlaid waveforms
  • Scripts ingested by drag and drop — Numbers, Excel or Word — with correctable column roles
  • A review report that starts with what needs a decision, plus color-coded Excel and CSV exports
  • Batch tracking: the project fixes the universe of lines, each incoming batch checks off what arrived
  • AsRec export: a copy of the client's own Excel with only the AsRec column changed, differences in red
  • Assembles the outgoing delivery replicating the original folder tree, with rename rules and a plan you approve before it runs
  • Runs entirely on the machine — no account, no API, no internet; client material never leaves the Mac

By the numbers

  • 1,000 take pairs cross-checked in 295 ms
  • Loudness measured at ~115× real time
  • A 1,000-line script indexed in 235 ms
  • Batch transcription at ~1.2 seconds per short take
  • Recurring cost: zero — open models, on-device
  • A 9,000-take delivery: roughly two people for three weeks by hand — about three hours of machine time here, line by line
  • Read the method: the seven-layer dubbing QC checklist →

Does this look like your bottleneck?

This system started as one specific problem in one real company. Tell us yours in a sentence — within 24 hours we'll tell you whether it's buildable and whether the math is likely to work.

Show us the problem →

See all 13 systems →