The Dubbing Glossary: 30 Terms Creators Actually Need
Dubbing brings its own vocabulary. Some terms come from traditional film and broadcast work. Others arrived with AI. Whether you are new to dubbing or scaling from a few videos to dozens, knowing these words saves time in conversations with translators, engineers, and your own team. More importantly, the vocabulary teaches you what can go wrong, when to expect it, and how to fix it before it ships.
How Dubbing Actually Works: The Vocabulary Map
Dubbing workflow has three stages, and each stage has its own vocabulary. Understanding what happens at each stage helps you predict problems and ask the right questions when something sounds off.
The first stage is extraction: your original video goes in, and software extracts the dialogue. This is where ASR (Automatic Speech Recognition, also called speech-to-text) pulls the words from the audio. ASR accuracy depends on audio quality, background noise, and whether it recognizes the speaker's accent. A clean recording with one speaker in a quiet room will convert to text in seconds. A noisy location shoot or heavy background music will confuse ASR and leave gaps.
The second stage is translation and synthesis. Your extracted text gets translated into the target language. Then TTS (Text to Speech) software reads it aloud. Modern neural TTS systems can vary pitch, pace, and emotion based on written instructions, so a translated phrase about good news can sound genuinely happy, not robotic. This is where prosody comes in, the rhythm, stress, and intonation that make speech sound natural. A sentence spoken with different prosody can sound angry, curious, sad, or neutral even if the words stay the same. Mediocre dubbing ignores prosody. Good dubbing preserves the original prosody or matches it to the translation.
The third stage is assembly and delivery: the dubbed audio gets synchronized with the original video, mixed with the original music and sound effects, and delivered in the formats your audience uses. This is where timing and loudness matter most.
Starting from the Audio: Foundations That Prevent Failure
Before dubbing even begins, the original recording has properties that will either help or haunt you. Sample rate is how many times per second audio is recorded, measured in Hz or kHz. Video typically uses 48 kHz. If you dub at a different sample rate and forget to convert, the audio will play too fast or too slow, creating sync problems that are hard to diagnose until the video is live.
Bit depth is how much information is stored for each audio sample. 16-bit is broadcast standard. 24-bit is higher quality but creates larger files. For dubbing, 16-bit is sufficient, and smaller files mean faster uploads and processing.
Your audio arrives as either mono or stereo. Mono is one audio channel. Stereo is two channels, usually left and right. Most dialogue is mono. Music and ambience are often stereo. Mixing them incorrectly creates phase issues, where frequencies cancel each other out and the result sounds hollow or thin.
Reference audio is a clean recording of a voice used to train or guide an AI voice cloning model. A good reference is 30-90 seconds, free from background noise, and covers a range of emotions or speaking styles. The better your reference, the more natural the cloned voice sounds in the target language.
Loudness and Dynamics: Why Your Dub Sounds Quiet
After your dub is synthesized, audio engineers measure loudness using LUFS (Loudness Units relative to Full Scale). This is not the same as raw volume. LUFS accounts for how human ears actually perceive sound, giving more weight to midrange frequencies that we hear more easily. YouTube recommends -14 LUFS. Broadcast television is -23 LUFS. Podcast platforms vary. Dubbed audio often comes out quieter than the original because the text-to-speech engine doesn't know how loud to speak. A final normalization step corrects this, bringing the dub to the target LUFS for its platform.
Headroom is the space between the loudest peak in your audio and the maximum level (0 dB). Leaving 3-6 dB of headroom prevents distortion when you normalize or mix. If you record too hot, every loud word will clip, creating digital distortion that cannot be removed.
Dynamics is the difference between the quietest and loudest parts of audio. A voiceover with flat dynamics sounds lifeless and like a robot. A good dub preserves or recreates the dynamic range of the original: loud moments feel powerful, soft moments feel intimate.
Compression is the tool that controls dynamics. It reduces volume when audio gets too loud, while making quiet parts audible without distortion. Light compression glues dialogue together, making it feel cohesive. Heavy compression flattens everything, removing the personality from the performance.
Timing and Synchronization: The Precision Problem
Timecode is a precise time label for every frame of video, written HH:MM:SS:FF (hours, minutes, seconds, frames). Timecode lets you mark exactly where dialogue starts and ends, helping you synchronize dubbed audio to video down to the frame. If your sync is off by even 100 milliseconds, viewers will notice the lag between speech and lip movement.
Sync means audio and video playing together at the right moment. Poor sync makes dialogue sound delayed or misaligned, and it destroys immersion faster than any other flaw. Dubbing requires pixel-perfect sync.
Latency is a delay between input and output. In recording, low latency means your voice appears in your headphones instantly so you can hear yourself perform. High latency makes recording feel sluggish and makes it hard to stay in time with music or background dialogue. AI dubbing systems have built-in latency in their synthesis, which is why timing often needs adjustment after generation.
Translation and Localization: When Words Break Timing
Dubbing is not translation plus audio. Translation affects timing directly. When you translate English into German, the German text gets longer. German and Russian expand significantly. Japanese and Chinese compress. If you record your dub with the exact same timing as the original English, an expanded translation will overrun the dialogue space, creating echo or overlap with the next speaker.
The solution is knowing expansion rates. Languages expand or compress differently: Germanic languages like German tend to be verbose, Romance languages like Spanish vary, and East Asian languages like Japanese often compress. By understanding how your target language behaves, you can cut scripts to fit the original pacing, omitting less important words in the translation to keep the rhythm.
Source language is the original language of the video. Target language is the language you are dubbing into. Spanish, German, and Mandarin are common targets for English content because they have large speaker bases.
A glossary is a list of terms and their approved translations. Music videos need glossaries to keep band names and song titles consistent. Medical videos need them so technical terms are translated the same way every time. Technical videos need them so product names and feature descriptions stay consistent across all language versions.
The End-to-End Workflow: Where Mistakes Happen
Here is a real workflow for dubbing a 5-minute educational video from English into Spanish:
Day 1: Extraction and setup. You upload the video. ASR runs and generates a transcript. You review the transcript for errors. ASR missed one phrase in a noisy section, so you correct it. You set up a glossary with product names and technical terms that should stay consistent in Spanish. You also split the audio into stems: dialogue, music, and effects. This gives you flexibility later if the effects stem needs adjustment independently.
Day 2: Translation and synthesis. A translator converts the script to Spanish, noting which passages expanded beyond the original timing. Spanish translations often run longer than the original English, so the translator shortens a few sentences to keep pacing tight. You provide a reference audio clip of a native Spanish speaker so the TTS engine can match the speaker's vocal characteristics. The system synthesizes the Spanish dialogue. You spot-check 10 clips by ear and find two that sound unnatural: one has the wrong emotion, one has odd pacing. You re-synthesize those two with adjusted prompts.
Day 3: Sync and mixing. The Spanish audio gets time-stretched to match the original pacing, preserving sync with the original video. This is where timecode matters: you verify frame-accurate timing at the start and end of each sentence. Music and sound effects are added back on their own stems (separate tracks) so they can be adjusted independently. You check LUFS levels on the dialogue stem and apply compression to even out the loudness. The dialogue sits at -14 LUFS now, matching the original. You verify that headroom is at least 3 dB, so no unexpected peaks will clip.
Day 4: QA and delivery. A second person reviews the video for mispronunciations, timing issues, and mixing problems. They spot one word that sounds like a different Spanish word, creating confusion. You re-synthesize that one phrase with a glossary entry to prevent future mispronounciation. Final delivery happens as both an MP4 (burned in as the default audio) and as a separate WAV stem for archival and future remixes. You also export an SRT subtitle file with timecodes so non-Spanish speakers can still follow the dialogue.
Common Failure Modes: Symptoms and Fixes
Symptom: Audio sounds muffled or quiet. This is usually a LUFS issue. Check the loudness of your final mix. If it is below -16 LUFS, normalize it upward. If it is above -12 LUFS, reduce it to avoid clipping on some platforms. Use a reference track from your source video to match perceived loudness. If the problem only affects certain words, you may have applied compression too aggressively. Reduce compression ratio and re-mix.
Symptom: Dialogue plays in sync but looks delayed on video. This usually means sample rate mismatch. Verify the sample rate of both the original audio and the dubbed audio. They must match exactly. If the original is 48 kHz and your dub is 44.1 kHz, convert one to match the other before syncing. Another cause is latency introduced during TTS synthesis. Check whether the dubbing platform added delay during generation and compensate in the sync alignment.
Symptom: Spanish or German translation runs over the original dialogue timing. You forgot to account for expansion. Calculate expansion rate for the language pair and cut the script accordingly. A common fix is omitting filler words (um, uh, so) and contracting phrases. Another approach is slowing playback slightly (90-95% speed) in the target language to fit the original timing without sounding rushed.
Symptom: Voice sounds robotic, no emotion or personality. TTS did not get enough guidance on prosody. Re-synthesize with more detailed emotional prompts (specify angry, calm, excited, sad rather than just reading neutral). Also check that your reference audio covers a range of emotions. A reference that is always neutral will produce a neutral dub regardless of prompts.
Symptom: One specific word is mispronounced. This is usually a TTS model limitation or a glossary error. If it is a proper noun or brand name, add it to the glossary with a phonetic spelling. Re-synthesize that phrase. If it is a common word, it may be a model issue, and you will need to record a human voiceover for that segment. Some platforms allow you to override specific phonemes, which is faster than re-recording.
Checklist: Quality Gates Before Publishing
| Issue | Check | Fix |
|---|---|---|
| Loudness | Is final mix at platform-correct LUFS? | Normalize dialogue stem to spec |
| Sample rate | Do all audio files use the same kHz? | Convert outliers before mixing |
| Sync drift | Does dialogue line up visually at the end? | Re-time-stretch if drift > 100ms |
| Translations exact | Do proper nouns match the glossary? | Re-synthesize mismatches |
| Emotional prosody | Does tone match the original performance? | Re-synthesize with emotion prompts |
| No clipping | Are there any distorted peaks at > -0.5 dB? | Reduce gain and re-normalize |
Frequently Asked Questions
Q: Do I need an M and E track? An M and E track is a version of the dubbed video with all dialogue removed, leaving only music, ambience, and sound effects. You only need one if you plan to recut the video later without losing the soundtrack. For a one-off upload to YouTube, you do not need it. For a distributed series where multiple editors may remix footage, an M and E track saves time.
Q: What format should I use for final delivery? For video platforms like YouTube, MP4 with burned-in audio is simplest. For professional workflows, export WAV stems (dialogue, music, effects) separately so future remixes are possible without re-dubbing. MP3 is sufficient for social media clips where file size matters more than archival quality.
Q: Should I use SRT or VTT subtitles? SRT is widely supported and human-readable. VTT is more feature-rich for web video, allowing inline styling. For YouTube and most platforms, SRT is sufficient. Use VTT only if you need font or color styling in your subtitles.
Q: What does burn down mean and when would I use it? Burn down is gradually fixing issues. In a dubbing project, you might burn down a list of five words that need re-recording, or fixes needed across multiple language versions. You prioritize the most noticeable issues first, test the fix, and repeat. This term comes from project management, but it applies to dubbing when quality issues accumulate during production.
What to Do Next: Build Your Glossary
You now know what to listen for and what to ask when something sounds wrong. The next step is practical:
-
Create a glossary. Write down every brand name, character name, and technical term in your video. Get the approved translation for each language you will dub. Add phonetic spelling if the TTS engine gets it wrong.
-
Save reference audio. Record 60 seconds of a native speaker in your target language saying a range of emotions: happy, sad, neutral, excited. Use this reference every time you synthesize to make the dub sound natural and consistent.
-
Check sample rates. Before you start, verify that your original video uses 48 kHz audio. If it uses something else, convert it. This prevents sync problems later.
-
Plan for expansion. Research expansion rates for your language pairs. Match your script cuts to the expansion you observe for each language. Test on the first minute to confirm.
-
Test one complete dub end to end. Do not dub your entire channel at once. Pick one video, dub it completely, get feedback, and apply those lessons to your next batch. This reveals workflow issues before they scale.
The Vocabulary is a Toolbox
These thirty terms are not just definitions. They are tools for diagnosing problems and building better dubs. Sync issues trace back to timecode or latency errors. Loudness issues point to headroom or LUFS problems. Stilted performance points to prosody. Once you know what to look for, you can prevent most problems from reaching your audience.
Dubbing at scale requires this vocabulary. Your translators need it to explain expansion. Your audio engineers need it to set gates and check levels. Your team needs it to talk about what they hear without confusion. Start with these thirty. Add more as your workflow deepens.
🚀 Start Dubbing Your Videos Today
DubLab uses AI to translate your videos into 92+ languages in minutes.