Identify Which Track Is Changing

Listen to the affected voice, music, and source separately. If speech alone is uneven, use the segment-volume guide. If music rises after each line and falls at the next, inspect recovery between windows.

A source-dialogue fragment emerging in a pause is a separate handoff problem requiring complete source sentence boundaries. Music-window merging cannot solve it alone. Distinguish a track’s own strong beats or arrangement changes from gain automation; constant gain does not guarantee constant perceived loudness.

Bridge Short Pauses and Decide Longer Breaks

The ordinary narration path sorts music-ducking windows and joins adjacent gaps strictly shorter than the merge threshold, using the lowest gain in a merged span. This prevents a brief rise followed immediately by another fall. Equality does not satisfy that less-than rule.

Ramp down before the window, hold attenuation, then recover after it. Assembly configuration defaults are 0.3 seconds for ramps and 1.5 seconds for merging. These are separate controls, not universal standards.

Preserve intervals intended for complete source dialogue, action sound, or meaningful silence. An excessive threshold may suppress a deliberate source block too. Source and music recover differently and need separate checks; the renderer does not interpret the creative allocation board automatically.

Worked Example: Holding Across a 0.4-Second Gap

Suppose a chosen 12-second paper timeline contains narration at 1–3, 3.4–5, and 8–10 seconds. Music gain is 0.18 normally and 0.10 under voice; ramps are 0.3 seconds and the merge threshold is 1.5. These are illustrative choices, not listened-to audio or universal recommendations.

IntervalMusic arrangement
0–0.7 sOrdinary gain 0.18.
0.7–1 sRamp to 0.10.
1–5 sMerge the first two windows: 0.4 is below 1.5; hold 0.10 across the gap.
5–5.3 sRecover to 0.18.
5.3–7.7 sRestore for this chosen longer break; do not add music or dialogue.
7.7–8 sRamp down before the third line.
8–10 sHold 0.10.
From 10 sRecover by 10.3, then hold 0.18. No final fade to zero is specified.

The three-second gap stays separate. Gains describe amplitude, not percentages of perceived loudness. Listen for musical accents masking words. Important source dialogue in 5–8 seconds requires its own actual arrangement; “restore” in a table cannot establish suitability.

This calculation explains merge and ramp times. Later mixing, normalization, and encoding mean these gains do not establish final loudness compliance.

Identify the Audio Path Before Changing Settings

Ordinary narration can duck source and handle optional music using these windows. source-mix creates no narration windows and does not run this narration ducking. Frozen audio copying does not remix.

The explicit full-mix path still named narration bypasses legacy ducking, ambient BGM, and automatic loudness operations. Revise the independently adopted bed and its arrangement at production; changing legacy ramp or merge settings cannot rebuild a curve on this path.

Keep the actual mode, voice placements, bed intervals, and affected times for review against real inputs. An old timing sheet cannot describe new voice gaps, and rerendering should not overwrite adopted sound that must remain.

Listen to the Complete Handoffs Before Reviewing Picture

Listen across a line’s ending, pause, and next opening. Check sudden rises, masked source dialogue, and lost intentional quiet. Then hear the whole passage; one smooth boundary does not establish the section’s rhythm.

For continued pumping, inspect actual windows and merge conditions. For continuously suppressed source, check accidentally bridged longer blocks. For one faint line, return to voice segments. For a strong music accent, locally change music or its selection rather than lowering the entire voice mix.

Review the current export and whether the editable timeline uses the same windows and curve. A written file, correct calculation, or overall loudness target cannot replace listening to this output.

FAQ

Will lowering all music fix it?

It may reduce masking without changing recovery at every pause. Separate excessive level, short-gap pumping, and musical accents before choosing overall gain or local windows.

Should every pause stay ducked?

Not always. Short pauses may bridge; complete source dialogue, action sound, longer breaks, and intentional silence need deliberate treatment. A gap threshold is not a voice-to-source quota.

Do it with the skill

Use video-assemble to investigate bed swells between voice lines. Include audio mode, actual placements, music and source roles, affected times, and listening observations. Separate music recovery, source-dialogue handoff, and uneven voice. Check short-gap merging and ramps while preserving intentional source blocks. For an explicit full mix, revise the adopted bed instead of expecting legacy ducking.

Read the method: Short-gap merging and ducking ramps · Separate source and music mixing paths · Ramp, gap-merge, and music-gain settings · Ordinary narration and explicit full-mix boundaries

About Video Recap SkillsThe video-assemble skill on GitHub

All guides