Assign duties to the three sound sources

Video Recap treats source sound, narration and music separately, with creative decisions preceding parameters. Record what viewers must hear in each passage. Meaningful silence is content, not an empty space awaiting music.

Original example: a character says, “You have the key,” opens an empty cabinet, then looks at another person. The line establishes possession, the opening sound completes the action, and brief narration explains why they sought the records. Loud music can mask both the line and action.

Listen before adding music. If source dialogue is already unclear, inspect the selected stream or original material instead of blaming BGM or inventing its words through narration.

Record local sound priorities

PassageRecord
Crucial dialogueMust remain clear: Complete key line Music: Lower or omit Source sound: Preserve sentence boundaries
Empty cabinetMust remain clear: Opening and pause Music: Quiet continuation or absence Source sound: Keep useful action sound
ContextMust remain clear: Connected narration Music: Lower beneath narration Source sound: Duck nonessential source sound
ResponseMust remain clear: Answer and hesitation Music: No sudden rise mid-answer Source sound: Restore at a reliable pause

This is an editorial choice, not a universal ratio. Fixed narration coverage or one volume difference cannot account for material, voices and frequency overlap. Listen to the actual result.

Complete Eight-Second Example: Keep the Line and the Reply

Continue the original key-and-empty-cabinet scene through a complete sound handoff. These times are selected paper windows; remeasure actual speech, action, and pauses in the source. If a whole sentence cannot fit, revise the window rather than removing its ending.

TimeSound and music
0–1.5 secondsA’s complete source line, “You have the key.” No additional BGM or competing narration.
1.5–3 secondsB hands A the key; A receives it and opens an empty cabinet. Keep its action sound and pause; do not use a musical hit to invent another person’s discovery.
3–6 secondsNarration: “They came for last night’s duty records.” Supply the established purpose, lower nonessential background, and add no BGM.
6–8 secondsA looks at B. B replies, “I didn’t take them.” Retain the complete voice and hesitation, without narration announcing who removed the records.

This candidate omits additional music for all eight seconds to recover speech and pauses. It does not require music-free videos. B’s reply is a denial, not proof of innocence; the empty cabinet does not establish that records were never placed there.

In the ordinary entry, removing BGM removes only the separately added score. If source music and speech are already mixed in one track, BGM_VOLUME cannot independently lower that embedded music. Find usable clean sound or explicitly revise the actual bed; do not claim the line was recovered merely by changing a setting.

For source-interval reconstruction, select corresponding real source ranges. source_score.py produces source, score, combined beds, and a receipt rather than a final video. The strict score:{"kind":"none"} form expresses no additional score; the full-sound path can then add selected narration. The producer does not choose dialogue for you.

If later music needs local lowering, the cited raw-score path has one continuous playhead, a whole-score gain, and entry/exit fades. It does not automatically duck each table row. Produce and review the required music track separately, then adopt it as a frozen stem with the required actual length and sound. Do not apply its completed gains and fades again. Finally listen through all eight seconds; this candidate was neither rendered nor heard for the article.

Confirm the assembly path before changing settings

In the cited ordinary assembly entry, BGM_PATH selects music; BGM_VOLUME and BGM_DUCKING_VOLUME control its levels. Source sound has separate idle and under-narration settings. Adjust the conflicting source and render a short candidate. Final loudness normalization does not replace local balance.

An explicitly adopted complete-bed mix bypasses legacy BGM, ducking and normalization settings. Changing those settings will not repair its music. Do not send a finished bed through legacy ducking again; revise the original intervals and music decision, then adopt the revised bed.

Parameters establish implementation, not intelligibility. Never remove necessary voice endings to satisfy a loudness target.

Once local music balance is repaired, apply the required whole-film loudness treatment and measure actual LUFS and true peak in the current encoded output. Short-drama audio craft warns that even two-pass loudnorm may miss the target; a target is not a measurement. Whole-film compliance still needs line-by-line listening and cannot establish that the key line is clear.

If narration varies between loud and quiet before music is added, inspect voice segments. TTS RMS records describe that WAV stage rather than finished-mix loudness. Distinguish these problems before raising all segments together and masking essential original dialogue again.

Listen through complete handoffs

Listen through crucial lines, narration boundaries, music entries and source restoration. Check for half-sentences suddenly becoming loud, music restarting without a story change and narration competing with essential dialogue.

Try less music or none. If two voices conflict, move narration or revise its text rather than only lowering music. Actual listening on speakers and headphones can reveal problems that numbers and reports cannot settle.

Watch the passage at normal speed and confirm that sound clarifies choices and results. An unreviewed mix remains a candidate, not proven clarity.

FAQ

Will lowering the whole video fix masking?

Usually not: both become quieter and their relative balance remains. Adjust individual tracks and, when needed, change sound duties or narration placement.

Do it with the skill

Ask video-recap to assign essential dialogue, narration and music roles, then use video-assemble on the corresponding path and provide a short listening candidate.

Read the method: Assembly sound roles and music settings · Audio modes and processing boundaries · Post-mix loudness and line review · Source intervals, no additional score, and frozen stems · Full-sound path for beds and selected narration

About Video Recap SkillsThe video-assemble skill on GitHub

All guides