Locate the Layer Where Speech Is Lost
Listen to the adopted video and identify the word, phrase, and time, then listen to its segment separately. Do not regenerate everything before locating the fault.
| Comparison | Next step |
|---|---|
| Approved text is complete; the TTS request is shorter | Check current text, cleanup, and shortening instead of using an old draft to describe this synthesis. |
| Request is complete; the WAV omits words | Repair the specified synthesis and listen again. |
| WAV is complete; the video omits it or its tail | Check consumed files, actual placement, mix, and fit results. |
| Caption is complete; speech is not | Displayed text proves no utterance. |
Video Recap records authored text, the TTS request, and the WAV path. spoken_text is requested-text evidence, not a transcription recovered from the WAV. Even a matching string still needs listening.
What Approved-Text Protection Actually Prevents
The cited ordinary narration path may shorten an overlong segment at a sentence boundary and mark truncated. Adopted text that must remain exact needs the explicit approved-text policy; asking “do not omit words” is not proof it was enabled.
--preserve-approved-text retains authored text and the cleaned request. A complete line exceeding its window and tempo budget blocks the run; required failure cannot become a deliverable partial result. A segment containing only removed stage cues also fails. An archived earlier success does not represent this failed run.
The policy protects input text and processing, not correct pronunciation, audible words, or an unobscured mix. Check cleanup, shortening, provider audio, and placement separately. Authorized rewriting remains local and preserves negation, quantities, and conditions.
Paper Example: One Missing Negative Reverses Meaning
This diagnostic exercise contains no generated or listened-to media. The approved line is “The light is still on; do not close the lid yet,” also shown in the caption.
| Assumed result | Check and repair |
|---|---|
| The request says only “Close the lid” | Text changed; restore the condition and negative, resolve the window, then synthesize. |
| Request is complete; assume the WAV says only “Close it” | Matching text is insufficient; remake the segment and listen for the negative. |
| Assume WAV is complete; the video loses its opening | Check placement and available room instead of rewriting everything. |
| File is missing or no safe placement exists | Unplaced is incomplete; supply valid audio or repair timing and inspect the new result. |
Record approved wording, actual request, audible speech, and status separately. An unheard result stays unchecked rather than being copied from the script into a listening field. Prioritize negatives, numbers, names, and ordering conditions.
Repair example: wait for the light, then close the lid
Keep the original line, “灯还亮着,先不要合盖” (“The light is still on; don’t close the lid yet”). Select an original eight-second miniature-model demonstration: at 0–4 seconds the indicator is on and a hand waits beside the open lid; at 4–6 it goes out and the teacher’s original audio says closing is now allowed; at 6–8 the lid actually closes. Narration uses only the original line at 0–4, leaving the later original speech. No source file or listening result exists here.
Assume approved copy and spoken_text contain the whole line, but the faulty WAV omits “不要,” reversing its instruction. Correct captions cannot repair that sound. Listen to the WAV, check that the finished candidate uses it and record the missing negation rather than assuming truncation or mixing before diagnosis.
Inspect caching before regeneration: This implementation reuses segments by text, settings and WAV identity. Rerunning unchanged input may reuse the faulty WAV. Preserve the affected audio and sidecar evidence, then invalidate only that segment and synthesize again. Confirm a new candidate was generated without clearing correct segments or treating old success metadata as current completion.
Enable approved-text protection and retain the original sentence. Actually hear the negation and final word in the new candidate. Continued omission remains a failure. If the complete voice will not fit 0–4, report the window conflict and revise the arrangement by author decision; do not cut the negation or cover the teacher at 4–6.
Once complete and placeable, let assembly consume the valid current audio and metadata, checking placement and captions. Play normally from lit indicator and waiting hand through extinguished light, permission and actual closing. Narration must not move permission earlier or cut the teacher’s ending. Record the version actually heard, whether the missing word returned and observations from the full demonstration.
Repair the Segment, Then Check Joins and the Full Video
Retain unaffected wording and audio, then confirm the repaired WAV is consumed by the current assembly. A changed duration needs checks of the next line, action, and captions. A cache is not proof of current use; synthesis is not placement.
Listen to the segment, its joins, and the full video at normal speed. Record whether the missing words returned and whether the tail, voice, or level now breaks. ASR can locate suspect wording but still requires listening, especially for disputed pronunciation.
This page addresses missing speech content. Source dialogue cut at an edit, incorrect captions, and music masking speech have separate diagnoses.
FAQ
Does a complete caption prove the voiceover is complete?
No. Captions may come from the draft rather than audio. Check the current segment and video, recording text matching separately from audible completeness.
Do I still listen after protecting the approved text?
Yes. Protection blocks processing changes; provider omissions, wrong files, placement, and mixing still need checks.
Do it with the skill
Ask video-recap to diagnose this specified voiceover or assembly problem. Identify current media and time, observations, missing evidence, and a limited repair. Preserve approved text and unaffected sound and picture; mark ungenerated or unheard parts unchecked rather than treating metadata as listening approval.
Read the method: Approved-text and partial-result boundaries · Protection policy and duration blocking · Requested text and segment records · Missing files, placement, and fit · Assembly roles for original speech and narration