Identify the actual timing mismatch
Video Recap synthesizes timed narration in segments. In cut mode, those times belong to the edited output, not the source clip. Check positions, overlaps and space reserved for essential original dialogue.
Original paper example: segment six occupies output seconds 18–22, a four-second window, while a candidate complete reading lasts six seconds. These numbers illustrate diagnosis rather than measured audio. Record authored text, spoken text, complete duration and end position; subtitle line count is not enough.
Fix incorrect timing first. For a correct but overloaded window, consider text compression, picture changes or placement rather than blaming provider speed.
Compress repetition while preserving cause
Overloaded draft: “Because he had previously promised to keep the key, it was only after discovering the empty cabinet that he finally realized he must immediately find the person who had given it to him.”
Revision: “He promised to keep the key. With the cabinet empty, he must find the person who handed it over.” Preserve promise, discovery and next goal while removing repeated time markers. Resynthesize and listen; word count cannot prove the four-second fit.
If retained original dialogue already establishes the promise, narration may shorten further. That depends on viewers actually hearing the premise, not on evidence removed from the edit.
Fitting Six Seconds into Four: Check What Comes Next
Continue the paper 18–22 window and assume essential original dialogue occupies 22–24. Six seconds of narration starting at 18 would end at 24 and mask that dialogue. Changing an end field from 22 to 24 does not resolve the conflict.
| Choice | Repair and check |
|---|---|
| Text may change | Remove repetition, synthesize the revised draft, and confirm complete speech fits 18–22; word count cannot prove a four-second result |
| Text and original speed are fixed; picture may extend | If actual usable continuous footage exists, retain two more seconds here: narration occupies 18–24 and original dialogue with its picture moves to 24–26; update later positions too |
| Text is fixed; picture cannot extend | Find sufficient space elsewhere or redesign the passage; otherwise leave placement unresolved rather than clip the ending or mask essential dialogue |
The second row is an editorial option for this assumed case, not an automatic voiceover capability. Confirm additional footage exists and does not fabricate a response, then remeasure duration, captions, and effects. Stretching a still image by two seconds is no automatic natural solution. If real audio or rendered durations differ, rearrange using actual values.
Text protection need not freeze every performance choice: a suitable intelligible reading can be tried. Once an original-speed version is adopted, arrange around that decision instead of hiding inadequate space with downstream default speed changes.
Protect adopted text explicitly
In the cited video-voiceover version, default overflow may shorten at sentence boundaries and record spoken_text/truncated. For fixed text, explicitly use --preserve-approved-text: it retains authored evidence and fails when the complete reading cannot fit instead of quietly delivering less.
Resolve failure through an editorial choice: a longer window, context moved elsewhere or an explicitly revised draft. A downstream original-speed adoption contract must not be overridden by an old fitting cache.
Voiceover generates supplied text; it does not choose shots, mix or render subtitles. After changing placement, check downstream audio and subtitle text so they do not use different versions.
Segment metadata fit_status: pending_assembly means placement still awaits assembly; audio_duration is synthesized duration. Verify actual placed start/end after assembly. A successfully generated WAV does not establish complete delivery of six seconds inside a four-second picture window.
Listen to the complete final word
Listen to sentence endings, breaths and the next entry. Check intelligibility after speeding, complete conclusions and protected source dialogue. A complete WAV clipped during placement is still a failure.
Do not accelerate every segment for uniformity. Repair the overloaded passages first. Extend picture where appropriate; otherwise change text or placement.
Review the output timeline for overlap, missing required segments and text matching actual speech. Duration checks locate problems; listening and viewing establish intelligibility and narrative completeness.
FAQ
Why does speech overflow when subtitles fit?
Subtitle length is not reading duration. Pace, pauses and voice affect timing. Compare complete generated audio with the output window.
Do it with the skill
Ask video-recap to preserve fixed text and use video-voiceover’s approved-text protection. Report authored and spoken text, full duration and window before proposing a repair.
Read the method: Segment voiceover and fixed-text protection · Assembly placement and complete endings · Generated duration and pending-placement metadata