Check Playback, Then Isolate the Track

A fast player, sped-up source sound, separately accelerated narration, or different TTS delivery can sound similar but need different repairs. Restore normal playback, isolate source and narration, and compare with prior versions.

A higher-sounding voice need not have one technical cause. Do not immediately change narrator or rewrite the script. Find the operation first. Picture can accelerate while sound is placed separately, or both can intentionally change together.

Distinguish Rate Reinterpretation From Tempo Processing

The official FFmpeg documentation describes asetrate as reinterpreting sample rate without changing PCM data, affecting speed and pitch. atempo processes audio tempo. These are different operations; ordinary format resampling should not all be called pitch shifting.

The cited Video Recap narration path uses atempo to avoid moving pitch with playback rate. Tempo processing still needs listening for emphasis, endings, and voice texture; no arbitrary factor is guaranteed natural.

Check TTS speed, global acceleration, and segment fitting separately. Successive tempo factors multiply; the final setting alone does not describe the actual path.

Paper Plan: Ten Seconds of Narration in Eight Seconds

The author chooses eight seconds of picture and the exact Chinese line “她把名单折了回去,门外响起两下敲门声。” (she folds the list back; two knocks sound outside). Assume an unaccelerated candidate lasts ten seconds and all words must remain. No audio is generated or heard here.

  1. Preserve the original candidate and eight-second picture, without another picture speed pass.
  2. Try narration tempo alone at 1.25: theoretical 10 / 1.25 = 8 seconds, not a measured output duration.
  3. This plan has no added TTS acceleration or segment fit. On the cited adoption path, independently select global_atempo=1.25, bounded_segment_fit=false, segment_tempo_max=1.0, and cumulative soft/hard limits of 1.25. The actual consumer accepts valid independently declared policies; ambient defaults do not choose one for you.
  4. Measure the result, check the full ending and consumed candidate, then listen to the export at normal speed. An overlong output cannot be trimmed or secretly sped up again.

If this delivery is unsuitable, the author can enlarge the picture window or authorize a line edit. Fixed picture, fixed words, and unacceptable speed leave the arrangement unresolved. Valid parameters do not establish natural delivery or actual fit.

Identify the Actual Audio Path

The ordinary narration path processes speed and fitting. Independently adopted policies follow the cited reference and consumer code, with conservative defaults of original speed and no segment fit. A default is not the only valid selection.

The explicit full-sound branch bypasses legacy speed/fitting. Changing the old narration_speed there will not accelerate narration; revisit selected inputs and placements. A frozen-audio picture revision should not secretly alter sound. Identify the branch before checking operation records and the export.

Review Duration, Words, and Sound Separately

ProblemFirst check
Player speed misleads comparisonRestore normal playback on the same device.
Faster tempo changes pitch unexpectedlyInspect the actual method and filters.
More acceleration than expectedCheck TTS, global, and segment operations.
Fitting drops the endingRestore complete input and revisit window/speed.
Adoption is treated as good soundHear the candidate, joins, and current export.

Retain text, the speed decision, actual processed file, and heard problems. Leave listening unresolved when absent; theoretical eight seconds and valid settings are not a listening pass.

FAQ

Does pitch-preserving acceleration guarantee natural speech?

No. Emphasis, pauses, endings, and delivery may still be unsuitable. Separating tempo and pitch is useful, but the actual candidate and export still need listening.

Must narration speed up with the picture?

No. Follow the sound plan. Separately placed narration can retain speed with a revised window. A chosen shared speed still needs complete words, cumulative factors, and actual alignment checks.

Do it with the skill

In the Video Recap flow, provide original audio, target window, selected speed, and the faulty export. Ask for diagnosis against the actual video-assemble audio path. Preserve other tracks and picture, add no hidden second speed pass or tail cut, then check consumed files, duration, and sound.

Read the method: Actual atempo processing and no-trim fitting · Independently selected tempo policy and consumption · Explicit full-sound bypasses legacy speed changes · Actual independent tempo-policy validation

About Video Recap SkillsThe video-assemble skill on GitHub

All guides