Listen to Voice Segments Alone First

Listen to neighboring WAVs in actual order before comparing the mix with music. The same voice can produce different levels with delivery, references, pause length, or synthesis results; a matching voice name is insufficient.

DifferenceFirst check
Voice files already alternate loud/quietCurrent WAV levels, peaks, settings, and intentional delivery
Voice alone is consistent, mixed speech is unclearMusic, source sound, and local timing/balance
One word is unusually harshTransient or performance, rather than indefinitely lowering the entire segment
Soft speech is intentionalIntelligibility and context; do not flatten every delivery

Save the current candidate and begin with the most noticeable segments. Repaired standalone sound must be heard in the mix too. Headphones and phone speakers test actual intelligibility; a target field cannot guarantee it.

Segment Normalization Has a Peak Constraint

The pinned video-voiceover defaults normalize each 16-bit PCM WAV to −20 dBFS RMS with a linear peak ceiling of 0.98. These configurable version defaults are not universal platform delivery requirements.

The implementation computes whole-block RMS and desired gain, caps that gain by the peak constraint, then records the output. It multiplies the whole segment by one factor, rather than compressing a single excessive transient. A peak-constrained gain may not reach the RMS target.

Segment fields rms_dbfs_before, rms_dbfs_after, and peak_after record this normalization stage. Null normalization records are no proof of compliance. Later tempo processing, mixing, or encoding means these values do not automatically describe the final file. Inspect adopted WAVs and settings from this run rather than an old cache/report.

Whole-block RMS includes pauses; long silence lowers the average. Equal RMS can still sound different with pauses, voices, and frequency content. Listen and judge a pause’s purpose before changing delivery or timing rather than deleting needed breaths to align numbers.

Original Calculation: Not Every Segment Can Reach the Target

These assumed values illustrate gain limits; none were synthesized or auditioned. Ignoring sample rounding, use −20 dBFS RMS as the worksheet target.

SegmentInput and expected result
A: quiet, low peakRMS −26 dBFS, peak 0.20; about +6 dB multiplies amplitude by 2, reaching −20 and peak about 0.40
B: quiet, high peakRMS −26, peak 0.75; multiplying by 2 would reach peak 1.50, over 0.98. Maximum gain is about 1.31, or +2.32 dB: RMS about −23.68, peak about 0.98
C: louder overallRMS −18, peak 0.50; about −2 dB multiplies amplitude by 0.79, reaching RMS −20 and peak about 0.40

Gain can bring A and C closer. B remains quieter because of its high peak, not a missing setting. Raising B to −20 would violate this example’s peak constraint. Listen for intended emphasis, a harsh transient, or a source problem. Use a separate audio tool if transient treatment is needed, or revise delivery and regenerate, then inspect the current sound. The gain guard is neither a compressor nor a true-peak measurement.

A remaining numerical difference need not force regeneration if connected listening is intelligible and the emphasis suits the line. The aim is natural clarity rather than identical report rows.

Measure and Listen Again After Mixing

Segment normalization does not mix music, duck source sound, or render captions. Once voice files suit their roles, arrange music, original dialogue, and narration, then inspect the finished mix.

Short-drama audio craft calls for measuring actual LUFS and true peak in the encoded output. Two-pass loudnorm does not guarantee hitting a target exactly: peak or dynamic-range constraints can cause dynamic processing instead of linear gain. −20 dBFS RMS, target LUFS, and measured encoded LUFS are different quantities. A WAV sample peak of 0.98 also does not prove compliant true peak in the encoded film.

Even after whole-film measurement passes, listen to connected lines: a single segment may lose words, music may rise mid-line, a join may cut breathing, or a whisper may become a shout. Repair local masking in the mix and performance problems in the relevant voice segment rather than repeatedly normalizing the entire film.

Retain before/after candidates, actual settings, and this run’s measurements. Mark unmeasured or unheard outputs awaiting measurement/listening instead of reporting a preset as a result.

FAQ

Will setting every segment to 100% make them equally loud?

No. Equal playback gain does not mean equal input levels. Compare actual WAVs and delivery, adjust gain within peak constraints, and listen in sequence.

Why can one line be quiet when whole-film LUFS meets the target?

Whole-film loudness cannot establish local balance. A line may be intentionally soft, lower-level, peak-constrained, or masked. Listen to it alone and in the mix, then repair the actual cause.

Do it with the skill

Use video-recap to compare uneven voiceover segments in isolation. Inspect current WAV RMS/peak and normalization settings, explain peak-constrained segments without repeatedly raising gain, and preserve intentional soft speech/emphasis. Remix repaired segments, measure encoded-export LUFS and true peak, and listen.

Read the method: Segment synthesis and output evidence · RMS normalization and peak gain guard · Pinned normalization defaults · Post-mix loudness and true-peak measurement

About Video Recap SkillsThe video-voiceover skill on GitHub

All guides