Separate wording from timing
| Situation | First repair |
|---|---|
| Voice says “keep” but caption says “delete” | Check adopted audio and the subtitle text source. |
| Words match but appear early | Check cue boundaries or track offset without substituting another sentence. |
| Caption shows original dialogue while narration is heard | Check source-dialogue versus narration attribution. |
Compare names, negation and ordering words carefully. Visible text is not automatically a caption for the current voice: titles, signs and source subtitles have separate roles. This page repairs captions intended to represent narration; see subtitle boundary checks for timing.
Identify what each text and audio record establishes
The voice module records authored_text, spoken_text and automatic shortening. The first retains the supplied author text; the second records words sent for synthesis. Neither replaces listening to what the provider actually spoke. Default over-window handling may shorten copy; --preserve-approved-text prevents automatic deletion and treats an overrun or required-segment failure as unfinished.
Confirm the current WAV is actually in the final edit before checking captions still based on old copy. A matching filename or present metadata does not establish adoption. If wording changes, the author selects the final version, then regenerates only affected audio and checks captions and joins against the adopted sound.
Complete example: an old caption reverses keep and delete
An original editing lesson’s approved narration says, “Keep this section; the next section can be deleted.” Its TTS spoken_text matches, and listening to the segment and current edit confirms it. An imported old cue instead says, “Delete this section; keep the next.” An 18–22-second window is a teaching arrangement, with actual boundaries still needing audio review.
| Material | Example result |
|---|---|
| Approved copy and spoken-text record | Keep the current section; delete the next. |
| Actual segment and adopted sound | The same words are heard; old audio was not substituted. |
| Displayed caption | Opposite old text causes a content error. |
Keep the current sound and replace the cue with the approved sentence, then verify its actual boundaries. Do not replace correct narration to accommodate old subtitles or merely move the incorrect sentence two seconds later.
For this repository’s independent track integration, edit the cue within the complete current track and retain every other caption needed. The integration applies to zero-start output media with adopted AAC audio in adopted-packet-copy mode. The track replaces all generated captions rather than automatically merging a one-cue patch. Do not assume the same integration applies to a newly mixed narration route.
After export, listen and inspect this cue’s first and last display, neighbouring joins and retained captions. This is a repair plan without generated TTS, track or video; actual adopted audio and the real track determine final check values.
Verify actual words and coverage after structural checks
The track loader preserves text and boundaries while checking current picture, adopted audio stream and duration. It does not recognise speech, correct words or prove synchronization. A structurally valid full track can still omit a required sentence.
Check separately which audio is adopted, which narration or source lines need captions, each cue’s current words and actual sound/display boundaries. Changed time or media requires preparing the affected track again; stale input cannot represent the current version. Even a one-line wording repair must retain other required captions when replacing a full track.
FAQ
Does another transcription guarantee correct captions?
No. Transcription is a candidate; compare names, negation and order against adopted audio, then verify wording and boundaries in the export.
Do it with the skill
At video-assemble, compare approved copy, spoken-text records, adopted audio and current captions. Repair the actual wording or audio fault before timing. Use full-track replacement only on its supported route and preserve other cues; validation does not replace listening.
Read the method: Author text, spoken text, approval protection and segment reuse · Recorded author-text and synthesis-text fields · Complete subtitle tracks, binding and content-check limits