Locate the wrong layer
The data contract distinguishes narration from original-dialogue gaps. Narration text follows the adopted voice track; dialogue subtitles follow what is actually audible in a gap. Text already burned into source frames is pixels, which a text-file edit cannot erase.
In an original example, a speaker says “Not Gu Yan—Gu Yuan.” ASR produces “Gu Yan—Gu Yuan,” losing the correction. Listen and check the approved character list for spelling; do not infer the intended name from the plot. Mark unclear speech for confirmation rather than guessing.
If the recording is wrong, changing the subtitle creates a mismatch. Decide about the audio first. The following workflow focuses on dialogue that is already spoken correctly.
Correct the source that wins
Dialogue-gap priority is: user_subtitles.json, supplied SRT/ASS, corrected original_subtitles.json, then ASR fallback. An older supplied file can still win even after the lower-priority correction is fixed.
| Source | First action | Check |
|---|---|---|
| Supplied text | Correct the adopted version | Coverage of retained dialogue |
| Reviewed text | Update audible utterances | Match to the current edit |
| ASR fallback | Prepare supported corrections | No coarse timing presented as precise |
State the edit’s coverage. If a supplied file contains only part of the dialogue, plan how remaining lines are retained; do not assume line-by-line merging with lower-priority files. Correct wording is separate from reviewing cue boundaries.
Keep the two clocks separate
This fictional timing example is not rendered media. A source utterance at 110–113 seconds appears at 8–11 seconds in the current cut. Unqualified source times can put the subtitle away from its speech.
In this pinned version, a bare user_subtitles.json array defaults to the OUTPUT clock. An envelope with timeline and lines can specify source or output. Supplied SRT/ASS defaults to source time; original_subtitles.json uses output time.
A source-clock correction can use:
{"timeline":"source","lines":[{"start":110,"end":113,"text":"Not Gu Yan—Gu Yuan."}]}
This shows file shape only. Obtain actual intervals from listening and the current edit map. Precise sources are clipped to covered dialogue-gap intervals and may split across boundaries. ASR midpoint estimates provide different evidence.
Review the actual new output
After rebuilding from current files, review names, negation, cue edges and returns from narration, then watch the actual full deliverable. Look for duplicate old lines, text left across cuts, inaudible dialogue or competing burned captions.
Burning dialogue also depends on the mask and explicit replacement request. A file’s presence does not prove it appears. Replacing source-burned errors requires a picture-edit plan; overlapping new text is not a repair by itself.
If the project uses an explicit subtitle_track.json, it completely replaces generated subtitles; it is not a one-cue patch. Carry all retained cues and bind current media. No rendering or listening result is reported here; valid files do not replace playback.
FAQ
Why does export still show old text after correction?
Check for a higher-priority supplied subtitle file and whether the current version was rebuilt. The old text may also be burned into source frames. Correct the relevant layer.
Do it with the skill
Use $video-recap to correct only these dialogue subtitles. Listen to adopted speech; check names, negation, winning subtitle source, clock and gap intervals. Preserve other copy and audio. Report actual display and playback checks after rebuilding, or mark them unperformed.
Read the method: Dialogue subtitle sources and clocks · Complete explicit subtitle-track replacement