Background Music and Dubbing: Keep the Mix Clean
When you dub a video, the background music is still there. If the original has a strong musical bed under dialogue, your new voice track has to compete with it, and the result sounds crowded and hard to understand. This guide walks you through the main strategies to handle music during dubbing, common mistakes to avoid, and how to know which approach fits your source material.
The core problem
Dubbed dialogue sits on top of everything else in the video. If the original mix has music playing at full volume while a character speaks, your new voice will either be buried or the overall mix will sound unbalanced. You have three paths forward: dub over the music as is, separate the music first, or add music back after a clean dub. Which one you choose depends on your video type and your tooling.
The reason this matters is simple. Listeners judge audio clarity within the first few seconds. If they cannot understand your character's words, they will mute or skip the video. A crowded mix where dialogue and music fight for attention creates that strain. The goal is to create space for the dubbed voice without erasing the emotional context that music provides.
Understanding your source material
Before choosing a strategy, listen to what you actually have. Play the original video and pay attention to three things: whether the music and dialogue are on the same audio track or separate ones, how loud the music is relative to speech, and whether the music is essential to the video's meaning.
If you can see separate audio tracks in your editor or on the video platform, the music is likely isolated. YouTube, Vimeo, and professional editing platforms often separate them. If you have only one audio file and cannot import separate stems, the music and dialogue are mixed together.
Next, listen for the music's volume during dialogue. Does it drop when someone speaks? Does the mix already have ducking applied? If the original mix is already balanced with dialogue clear on top, you are starting from a better place. If the music is present at full strength while people talk, you will need to dull it down.
Finally, ask whether the music is background (ambient, transitional, filler) or foreground (critical to a scene's meaning, part of a performance). A cooking show with light background music needs different handling than a live concert or a music video. This distinction shapes your strategy.
Ducking: the fast path
Ducking means lowering the volume of the music whenever there is speech. It is simple, fast, and works for most videos.
Workflow for ducking:
- Import your video into a video editor (DaVinci Resolve, Adobe Premiere, or any timeline editor works).
- If the music and dialogue are on separate tracks, place them on different audio tracks in your timeline.
- On the music track, create volume keyframes (automation points) at the moment dialogue starts and where it ends.
- Lower the volume envelope downward during all dialogue passages. Start with 8 dB of reduction and listen.
- If the dubbed voice is still hard to hear, increase the reduction to 10 or 12 dB.
- In pauses between sentences (a beat, a breath), raise the music back to full volume. This keeps the bed present and prevents the mix from feeling hollow.
- Export a test file and listen on headphones and a phone speaker to verify the balance.
If the music is already mixed with dialogue in a single track with no way to separate them, ducking is your only option. You cannot reliably reduce just the music without affecting the original dialogue still embedded in the same file, so you dub directly over it with aggressive ducking applied.
When ducking works well:
Ducking is effective for voiceover-heavy content like tutorials, explainers, interview shows, and educational videos. It is also fast: a 5-minute video typically takes 20 to 30 minutes to duck properly. The tradeoff is that the music ends up permanently quieter, which is fine when music is not the focus.
When ducking falls short:
Ducking struggles with videos where music is as important as dialogue, such as music videos, live performances, or emotional drama scenes. Ducking these means choosing: either dull the music so much that it loses impact, or let it stay prominent and risk burying the dialogue. In these cases, a different approach usually serves better.
Stem separation: theory vs. practical reality
Stem separation is the idea that you can extract individual instruments or vocals from a mixed audio file. Tools like Demucs use machine learning to attempt this. The appeal is obvious: remove the original voice and background sounds, keep the music, and lay your new dubbed voice on top. A perfect stem separation would give you surgical control.
Here is what actually happens in practice. Stem separation works well on clean, professional studio recordings with clear instrument separation. Most online videos are not this. YouTube audio is compressed. Mobile phone recordings have background noise. Voiceovers are mixed at different levels than music beds. Even excellent separation tools will leak some of the original dialogue into the music stem and some music tones into the voice stem.
The result is artifacts that you will hear on careful listening: a ghost of the original speaker's voice in the music track, frequency holes where speech was removed, or a thinned-out, unnatural quality to the music. These artifacts are often worse than simply ducking the original mix.
How to test stem separation:
If you decide to try it, do this test first. Export the music stem from your separation tool and listen to it on its own (no dialogue dubbed over it yet). Listen for:
- Echoes or faint traces of the original speaker
- Sudden drops or silence where speech was
- Unusual frequency distortions or ringing tones
- Loss of bass or presence in the music
If you hear any of these, stop. A failed stem separation is not worth trying to clean up. The effort to manually edit out artifacts will exceed the time it takes to duck the original mix or use one of the other strategies.
If the music stem sounds clean, proceed: mute the original audio track, dub your new dialogue in silence or lightly over the music stem, then layer the extracted music back in at the end. Balance levels as described below.
Re-add music after the dub: the clean path
The cleanest approach is to remove all audio from the original video, dub your new dialogue into that silent space, and then mix the music back in fresh.
Prerequisites:
You need the music as a separate stem (a "music-only" audio file) from your source. This is available if:
- The video was professionally produced and the source files include a music-only track.
- The platform (YouTube, Vimeo) allows you to download or export stems separately.
- Your source video's audio was mixed with clear separation and you have access to the original project files.
If none of these apply, you cannot reliably extract music after dubbing, so this path does not fit.
Workflow for re-adding music:
- Export the original video as a picture-only file (video without audio). Most editors have an option to export video-only or mute all audio before export.
- Import that picture-only file into a clean timeline.
- Add your DubLab-dubbed dialogue as a new audio track below or above the video.
- Import the music-only stem as a separate audio track.
- Adjust the levels: dialogue should peak around -12 to -6 dB, and music should sit 6 to 12 dB below that.
- In sections with no dialogue (transitions, montages, intros, outros), raise the music back to -12 to -6 dB so the track feels complete.
- Export and listen for balance on multiple playback systems.
This path gives you complete control and produces the cleanest result, but it requires access to the music stem and more editing work.
Audio quality and compression: why it matters
The quality of your source audio affects which strategy will work. Professional videos with high-quality, lossless audio allow more aggressive stem separation and finer ducking control. Compressed audio from streaming platforms (YouTube, TikTok, Instagram Reels) limits your options.
When audio is compressed (MP3, AAC, or heavily compressed streaming formats), the frequencies overlap more. The music and dialogue are already less distinct, so stem separation becomes less reliable. Artifacts are more likely. In these cases, ducking the compressed mix as a whole is usually more effective than trying to separate stems.
Check your source format before choosing a strategy. If you are pulling audio from YouTube or a social platform, assume it is compressed and plan to either duck it or find a higher-quality source. If the original was professional production material, stem separation is more likely to succeed.
Decision framework: which path to choose
| Your situation | Best approach | Why |
|---|---|---|
| One audio file, music and dialogue mixed | Dub with ducking | No separation possible; ducking is fast and effective |
| Separate music track, professional mix | Dub with ducking | Quick, preserves the original mix quality |
| Music stem available, clean original | Re-add music after dub | Gives cleanest result and most control |
| Compressed online video, urgent deadline | Dub with ducking | Stem separation likely to fail; ducking is reliable |
| Music is prominent (concert, performance) | Re-add music or consider subtitles | Full ducking kills the music; alternatives may serve better |
| Very loud background noise + music | Dub with ducking | Noise makes stem separation unreliable |
The mix itself
Once your dialogue and music are in place, balance their levels carefully. This is where clarity and emotional impact meet.
Level balancing:
Dialogue should peak around -12 to -6 dB on your mix. Music should sit 6 to 12 dB below peak dialogue volume, typically between -18 and -12 dB. If the dialogue is constant throughout, music should stay recessed. If there are long stretches without dialogue (a scene transition, a montage, an establishing shot), raise the music back to -12 to -6 dB so the video does not sound hollow or abandoned.
Use a loudness meter to check your overall mix. Dubbed videos often come out quieter than originals because the speech is typically recorded at lower levels than a professionally mixed broadcast signal. Compare the loudness of music-only sections to dialogue-plus-music sections and verify they feel balanced, not jarring.
Common mistake: over-correcting for dubbed speech
Dubbed voice acting is often recorded in a controlled studio environment, usually quieter than broadcast dialogue. Do not compensate by cranking up the dubbed audio to match the original. Instead, lower the music slightly more and let the dubbed voice sit at natural levels. This prevents the dialogue from sounding loud and unnatural against the music bed.
Common mistake: leaving the music too low
If you reduce the music too aggressively, the video loses emotional context. Music cues emotion, paces action, and fills silence. A mix where music barely peeks out from under dialogue feels hollow. Aim for a balance where dialogue is always intelligible but music is present enough to feel intentional.
Testing on real playback systems
Do not rely on monitor speakers or studio headphones alone. Export a test version and listen on at least three systems: good headphones, a car speaker, and a phone speaker. Each reveals different issues.
Headphones show you fine details: phase issues, frequency imbalances, subtle artifacts. A car or phone speaker shows you how most viewers will actually hear it. If the dubbed voice feels buried in the car, you need more ducking, a lower music level, or a quieter original mix stem.
Record a note on what sounds wrong so you remember what to fix. "Voice buried when music plays" means increase ducking or lower music level. "Music disappears entirely" means you over-ducked and should reduce the ducking amount. "Everything sounds muffled" means your dubbed audio may be the problem, not the music.
Listen to at least two separate scenes: one with continuous dialogue and music, and one with pauses between lines. The pauses reveal whether your ducking/mixing feels natural or abrupt.
Common mistakes and fixes
Mistake 1: Ducking too early or too late
Dialogue and music ducking should synchronize exactly. If you lower the music a half-second before the speaker starts, or keep it low after they finish, the mix feels loose or sloppy.
Fix: Place your ducking keyframes right at the moment the dubbed audio starts and ends. Use your editor's waveform view to see the exact frame.
Mistake 2: Forgetting to raise music in silent gaps
If you dub a scene with pauses between lines, duck the music during dialogue but raise it back during the pauses. Forgetting this creates an eerie silence that feels broken.
Fix: After you finish ducking, do a second pass to identify every pause longer than a half-second and raise the music back.
Mistake 3: Using the original audio as a reference when it already has problems
If the original mix is already muddy or has the dialogue buried, copying its balance will perpetuate the problem. Do not assume the original mix is "correct" just because it is official.
Fix: Listen to your dubbed version with fresh ears against a reference track from a different video in the same style. Does it sound balanced compared to that standard, or does it sound like the original's problems?
What to do next
-
Identify what you have. Download or export your source video and check whether the music and dialogue are on separate tracks. If they are separate, note this in your project.
-
Listen and decide. Play the original mix at your actual playback volume (not monitoring level). Decide whether the music is background filler or emotionally important. This choice determines your strategy.
-
Choose your approach. Use the decision framework above to pick ducking, stem separation, or re-add-music. Start with ducking if you are unsure: it is fast, reliable, and works for most content.
-
Dub or re-dub. If you need a fresh dub with different background conditions, use DubLab to generate a new version with your chosen strategy in mind. If you already have a dub, move to the next step.
-
Test before exporting final. Always export a test version at lower quality first. Listen on headphones and a phone speaker. If the mix is off, adjust ducking or levels and re-export a new test.
-
Export your final mix. Once you are satisfied, export at your target quality and platform specifications. Include the exact ducking or level settings you used in your project notes so you can repeat them if needed.
🚀 Start Dubbing Your Videos Today
DubLab uses AI to translate your videos into 92+ languages in minutes.