AI Video & Voice Production: ElevenLabs, Veo, Runway and More · Editing and audio post-production · lesson 12 of 16 · 8 min
Audio post: AI music beds, sound effects and a clean mix
Why audio post decides perceived quality
Viewers forgive average visuals far more readily than bad audio. A clean voice at a consistent level, music that supports rather than competes, and a few well-placed sound effects make even simple slide-based videos feel professional. AI now covers each layer: voice (text-to-speech), sound effects from text prompts, music generation, and voice isolation to clean recorded speech. Your job is to mix them so speech always wins.
The layers of a mix
| Layer | Role | Typical level relative to voice | |---|---|---| | Voice | Carries the message | Reference (loudest, clear) | | Music bed | Emotion and pacing | Well below voice; duck further under speech | | Sound effects | Emphasis, transitions, realism | Short and purposeful, below voice | | Room tone / ambience | Removes "dead air" feel | Very low |
Ducking means automatically lowering music whenever speech plays. Most editors offer it (sometimes called "auto-ducking" or sidechain compression). Use it; it is the single biggest improvement to most amateur mixes.
Loudness: deliver consistently
Platforms normalize loudness, so mixing "as loud as possible" does not help and often hurts. Use a loudness meter (most editors include one) and aim for a consistent target across your channel. A widely used delivery target for online video is around -14 LUFS integrated with true peaks below about -1 dBTP; broadcast and some platforms specify other targets, so check each platform's current guidance and your client's spec. Consistency between your own videos matters more than hitting an exact number.
AI tools for each layer (ElevenLabs example)
- Sound effects:
client.text_to_sound_effects.convert(text=..., duration_seconds=...)generates effects from a description (between 0.5 and 30 seconds at the time of writing). - Music:
client.music.compose(prompt=..., music_length_ms=..., force_instrumental=True)creates a track from a prompt; use instrumental beds under narration. - Voice isolation:
client.audio_isolation.convert(...)removes background noise from recorded speech (Descript's Studio Sound does a similar job inside the editor).
Check each tool's terms for commercial use of generated music and effects on your plan, and keep a record of prompts and dates, as with generated video.
Hands-on: generate a music bed and two effects
import os
from elevenlabs.client import ElevenLabs
client = ElevenLabs(api_key=os.environ["ELEVENLABS_API_KEY"])
def save(stream, path):
with open(path, "wb") as f:
for chunk in stream:
f.write(chunk)
try:
bed = client.music.compose(
prompt=("Calm, modern instrumental bed for an educational explainer, soft piano and "
"light pads, 90 bpm, no vocals, no dramatic drops"),
music_length_ms=60_000,
force_instrumental=True,
)
save(bed, "bed_60s.mp3")
for name, desc in {
"whoosh": "Short soft whoosh for a slide transition, clean, no reverb tail",
"chime": "Gentle single notification chime, warm, friendly, not shrill",
}.items():
sfx = client.text_to_sound_effects.convert(text=desc, duration_seconds=1.5)
save(sfx, f"sfx_{name}.mp3")
print("done")
except Exception as err:
print("generation failed:", err)
Mixing recipe (any editor)
- Place voice first. Normalize so speech sits consistently; remove clicks and breaths that distract (keep natural breaths).
- Add the music bed. Lower it until the voice is effortless to follow, then enable ducking under speech.
- Add effects only where they mean something: a transition, a notification on screen, an emphasis beat. Fewer is better.
- Add very low room tone under silent visual moments so they do not feel like dropouts.
- Measure loudness on the full mix and adjust the master to your target.
- Listen on a phone speaker and on earbuds. Most viewers will hear it that way.
Worked example: a lecture series
An online academy voices lectures with a designed narrator. Early episodes used a stock music bed at the same level throughout, and learners complained it was distracting during dense explanations. The team switched to a gentle instrumental bed with ducking, reserved sound effects for scene transitions only, and standardized loudness across episodes. Complaints stopped, and editors now reuse one mix template for every lecture.
Second worked example: a 15-second product ad
A Dubai perfume retailer's vertical ad uses a generated instrumental bed, a soft "spritz" effect on the product shot and an Arabic voice line. The first mix had music and voice fighting in the same frequency range. Lowering the bed, ducking it under the line and removing a busy percussion layer made the voice clear on phone speakers. The team logged the music prompt, model and date in the production record and confirmed commercial-use rights for generated music on their plan before running paid media.
Pitfalls
- Music at a constant level under speech.
- Effects on every cut, which feels cheap and tiring.
- Mixing on studio headphones only; phones reveal problems.
- Using generated or library music in ads without checking commercial-use terms.
Video lecture: Audio post: AI music beds, sound effects and a clean mix
Lecture coming soon · 14 chapters · about 8 minutes. Read the full transcript below.
- Audio post: music, effects and mix
- Why it matters
- The restaurant analogy
- Four layers and one technique
- Loudness
- AI for each layer
- Example 1: a lecture series
- Example 2: a perfume ad
- Cleaning real recordings
- Watch me do it, part 1
- Watch me do it, part 2
- Measuring audio quality
- Common mistakes
- Recap and try this now
Lecture transcript
Audio post: music, effects and mix
Try this experiment. Take any video you like and turn the picture off. If you can still follow it and it still feels polished, the audio is doing its job. Now try the reverse: keep the picture and make the audio slightly muddy. Most people click away within seconds. In this lecture you'll learn the layers of a professional mix, how to use AI for music beds, sound effects and voice cleanup, how to deliver consistent loudness, and a mixing recipe you can reuse on every video.
Why it matters
Why does audio matter so much? Because it's how viewers judge quality before they judge content. Viewers forgive average visuals far more readily than bad sound. A clean voice at a steady level, music that supports rather than competes, and a few well-placed effects make even simple slide-based videos feel professional. And AI now covers every layer. The skill isn't generating sounds. It's mixing them so speech always wins.
The restaurant analogy
Here's an analogy. A mix is like a dinner conversation in a restaurant. The voice is the person talking to you. The music is the background playlist. The sound effects are the clink of a glass at the right moment. And room tone is the gentle hum that tells you you're in a real place. When the playlist is too loud, you lean in and strain to hear your friend. That's exactly what viewers feel when music competes with narration. The playlist should be felt, not fought.
Four layers and one technique
So there are four layers. Voice carries the message and is your reference, the loudest and clearest element. The music bed sets emotion and pacing, sitting well below the voice. Sound effects add emphasis, transitions and realism, short and purposeful. And room tone, or ambience, sits very low to remove the dead-air feeling in silent moments. And there's one technique to learn first: ducking. That means music automatically dips whenever speech plays. Most editors have it built in. It's the single biggest improvement to most amateur mixes.
Loudness
Now loudness. Platforms normalize volume, so making your mix as loud as possible doesn't help, and often hurts because it squashes dynamics. Use the loudness meter in your editor. A widely used delivery target for online video is around minus fourteen LUFS integrated, with peaks below about minus one decibel true peak. Broadcast and some platforms specify other targets, so check current guidance and any client spec. But here's the key idea: consistency across your own videos matters more than hitting an exact number.
AI for each layer
Where does AI fit? For sound effects, ElevenLabs can generate an effect from a text description, up to thirty seconds at the time of writing. For music, it can compose a track from a prompt, and you can force it to be instrumental, which is what you want under narration. For recorded speech, voice isolation removes background noise; Descript's Studio Sound does a similar job inside the editor. As with generated video, check commercial-use terms for generated music and effects on your plan, and log prompts and dates.
Example 1: a lecture series
First example, a simple one. An online academy voices its lectures with a designed narrator. Early episodes used a stock music bed at the same level throughout, and learners said it was distracting during dense explanations. The team made three changes. A gentler instrumental bed with ducking. Sound effects only at scene transitions. And standardized loudness across episodes. Complaints stopped, and editors now reuse one mix template for every lecture. Nothing fancy, just the basics applied consistently.
Example 2: a perfume ad
Second example, a business case. A Dubai perfume retailer's fifteen-second vertical ad uses a generated instrumental bed, a soft spritz effect on the product shot, and an Arabic voice line. The first mix had music and voice fighting, so the line was hard to hear on phone speakers. The fix: lower the bed, duck it under the line, and remove a busy percussion layer that clashed with the voice. The team logged the music prompt, model and date, and confirmed commercial-use rights for generated music on their plan before spending on ads.
Cleaning real recordings
Sometimes you're not generating the voice at all, you're cleaning a real recording. A founder records a message on a laptop in a busy office, with a fan humming and a door closing. Voice isolation, whether it's the ElevenLabs tool or Descript's Studio Sound, can remove much of that background. But don't overdo it. Aggressive cleanup can make a voice sound thin or robotic. Apply it, then compare with the original at the same volume. Often a moderate setting plus a quieter room next time is the best answer.
Watch me do it, part 1
Watch me do it. In the lesson's script, I create a client with my key from the environment. I compose a sixty-second music bed with a careful prompt: calm, modern instrumental for an educational explainer, soft piano and light pads, ninety beats per minute, no vocals, no dramatic drops. I set force instrumental to true. Then I loop over two effect descriptions, a short soft whoosh for slide transitions and a gentle notification chime, each one and a half seconds long. Everything saves to files, and any error is caught and printed.
Watch me do it, part 2
Now the mix, in any editor. Voice first, leveled so speech sits consistently, keeping natural breaths. Then the music bed, lowered until the voice is effortless to follow, with ducking on. Effects only where they mean something, like a transition or a notification on screen. A very low room tone under silent visual moments. Then I measure the loudness of the full mix and adjust the master to my target. Finally, the step people skip: I listen on a phone speaker and on earbuds, because that's how most viewers will hear it.
Measuring audio quality
How do you measure success in audio? Three signals. First, loudness consistency: every episode lands within a small range of your target. Second, audience signals: comments about sound, and whether viewers drop off in music-heavy segments. Third, rework: how often a video goes back because of an audio complaint from a client or reviewer. If all three are quiet, your template is working. If one flares up, listen to that video on a phone before changing anything else.
Common mistakes
Common mistakes. Music at a constant level under speech. Effects on every cut, which feels cheap and tiring. Mixing on studio headphones only, when phones reveal problems you'd never hear otherwise. Chasing loudness until everything is squashed. And using generated or library music in ads without checking commercial-use terms. Notice that none of these are technical failures of AI. They're mixing decisions, and a template plus a phone check prevents almost all of them.
Recap and try this now
Recap. A mix has four layers: voice, music bed, effects and room tone, and speech always wins. Use ducking. Deliver consistent loudness, commonly around minus fourteen LUFS for online video, after checking platform and client specs. Use AI for effects, instrumental beds and voice cleanup, with licenses checked and prompts logged. And test on a phone. Try this now: remix one of your existing videos with a ducked bed, effects only at transitions, and a loudness check, then compare the two versions on a phone speaker.
Key takeaways
- A mix has four layers (voice, music bed, effects, room tone) and speech must always win.
- Ducking music under speech is the biggest single improvement to most mixes.
- Deliver consistent loudness (commonly around -14 LUFS for online video) after checking platform and client specs.
- Use AI for effects, instrumental beds and voice cleanup, but check commercial-use terms and log prompts.
Try it
Remix one existing video with a ducked instrumental bed, effects only at transitions and a loudness check, then compare old and new versions on a phone speaker.