Freeze the Final Output, Then Record at Least Three Checkpoints

In ZenStory's video-recap workflow, subtitle times use the final output clock. If cues were first timed to source media, establish a source-to-edit mapping before placing every cue boundary on the edited output clock. After that mapping, the creator still replays the video picture intended for release plus the narration actually adopted in that output and confirms the speech and cue starts and ends.

Label the exact video and narration versions. Choose one line with a clear spoken start and end near the beginning, one near the middle, and one near the end. If a line contains overlapping speech, an indistinct onset, or an unclear ending, choose a clearer line. Every time in the log must come from playback observation; write “unknown” when the evidence is unclear instead of guessing.

EntryRecord
Near beginningLine marker: [verbatim line or unique marker]; Speech start: [observed/unknown]; Speech end: [observed/unknown]; Cue display start: [observed/unknown]; Cue display end: [observed/unknown]; Start error (subtitle − speech): [positive=appears late; negative=appears early]; End error (subtitle − speech): [positive=disappears late; negative=disappears early]; Note: [overlap, unclear ending, etc.]
Near middleLine marker: [verbatim line or unique marker]; Speech start: [observed/unknown]; Speech end: [observed/unknown]; Cue display start: [observed/unknown]; Cue display end: [observed/unknown]; Start error (subtitle − speech): [same sign rule]; End error (subtitle − speech): [same sign rule]; Note: [note]
Near endLine marker: [verbatim line or unique marker]; Speech start: [observed/unknown]; Speech end: [observed/unknown]; Cue display start: [observed/unknown]; Cue display end: [observed/unknown]; Start error (subtitle − speech): [same sign rule]; End error (subtitle − speech): [same sign rule]; Note: [note]

A start error supports a decision about the start boundary only; an end error supports a decision about the end boundary only. When only a cue's start is observable, keep its end marked “unknown” and measure it before deciding whether to edit that boundary. The beginning, middle, and end form an initial pattern check. Add checkpoints around transitions, sections, or local recuts where another problem appears.

This is an editorial observation card, not the native subtitle_track.json format. The native track uses the final output clock and current media bindings. Its loader preserves authored text and boundaries; it does not align speech, map source time, or repair cues. Record every field separately in the two-column card rather than replacing them with “synchronized.”

Choose the Repair from the Error Pattern

This manual method compares observed start and end errors in the final media, then chooses a repair scope supported by that evidence. The subtitle track records boundaries on the final output clock; the creator confirms how those boundaries relate to actual speech by replaying the render.

EntryRecord
Start and end errors near the beginning, middle, and end are early or late by roughly the same amountDiagnosis: Global offset: the full set of cue boundaries is displaced while cue durations remain broadly intact; Evidence-supported repair: Shift the complete track by the verified common error: move late cues earlier and early cues later while preserving each cue's duration. This diagnosis requires matching evidence at both start and end boundaries
The beginning is close, the middle is farther off, and the end is farther still; start or end errors change steadily through the runtimeDiagnosis: Growing drift: the captions and final voice track do not share the same runtime change; Evidence-supported repair: First confirm that the SRT, narration, and video belong to the same final version; rebuild start and end times against the final narration, using verified segment anchors when needed. One global shift corrects only one position
Most checkpoints agree, but one cue or a small neighboring group has a bad start or endDiagnosis: Local boundary error: a local boundary or edit needs correction; Evidence-supported repair: Change only the start or end boundary supported by observed evidence, then replay that cue and its neighboring handoffs. Keep the unobserved boundary unchanged until evidence is available

A track can contain a global offset plus a few local errors. Fix the shared offset first, then re-observe subtitle start, subtitle end, speech start, and speech end for every checkpoint in the new export. Apply local retiming when that playback still shows a boundary error. If an end remains unknown, leave that boundary unchanged and measure it before editing. When evidence is sparse or contradictory, keep the classification at “insufficient evidence” and add observed checkpoints.

AI Can Organize Observations, Not Prove Synchronization from Text

A model can arrange the errors you recorded, identify a possible shape, and list positions for another review pass. Without playing the final media, it cannot know the true start or end points and must not invent timestamps. Copy the template below, leaving “unknown” wherever you lack an observed value.

I am checking subtitle sync against the final video and the narration actually used. Use only the manual playback observations below. Do not fill, infer, or assume missing times.

Final video version: [file/version]

Adopted narration version: [file/version]

SRT version: [file/version]

Checkpoints:

- Beginning: [line marker]; speech start [verified/unknown]; speech end [verified/unknown]; cue display start [verified/unknown]; cue display end [verified/unknown]

- Middle: [line marker]; speech start [verified/unknown]; speech end [verified/unknown]; cue display start [verified/unknown]; cue display end [verified/unknown]

- End: [line marker]; speech start [verified/unknown]; speech end [verified/unknown]; cue display start [verified/unknown]; cue display end [verified/unknown]

- Additional anomaly: [line marker and four verified times; mark every unobserved field unknown]

Calculate an error only when both values for that boundary are available: start error = cue start − speech start; end error = cue end − speech end. Classify only the available evidence as global offset, growing drift, local boundary error, mixed problem, or insufficient evidence. Preserve every unknown. Explain the evidence, but do not claim the track is synchronized. If evidence is insufficient, list the exact boundaries I must measure in the rendered video. Finally, propose only a correction scope supported by the evidence and list the start and end checkpoints that require manual replay after a new export.

The model's classification remains a working recommendation. The deciding evidence is playback of the corrected render, not the table or prompt itself.

Original Example: A Shared 0.8-Second Delay Plus One Local Error

The following original fictional project log illustrates the arithmetic and repair order; it does not report a media test. A five-minute paper-lantern folding tutorial uses the final narration version lantern-fold-final-03. Playback of the final output produced these observations:

EntryRecord
BeginningLine marker: “Fold the square paper corner to corner”; Speech start: 00:12.4; Speech end: 00:15.6; Cue start: 00:13.2; Cue end: 00:16.4; Start error: +0.8 s; End error: +0.8 s
MiddleLine marker: “Open the paper pocket, then press the crease flat”; Speech start: 02:18.1; Speech end: 02:21.7; Cue start: 02:18.9; Cue end: 02:22.5; Start error: +0.8 s; End error: +0.8 s
EndLine marker: “Thread the tassel through the lower paper loop”; Speech start: 04:41.6; Speech end: 04:45.0; Cue start: 04:42.4; Cue end: 04:45.8; Start error: +0.8 s; End error: +0.8 s
Additional anomalyLine marker: “Turn the folded point toward the center”; Speech start: 03:06.0; Speech end: 03:09.4; Cue start: 03:08.0; Cue end: 03:10.2; Start error: +2.0 s; End error: +0.8 s

The beginning, middle, and end are consistently +0.8 s late at both boundaries. That two-boundary evidence supports shifting the complete track 0.8 seconds earlier while preserving each cue's display duration. The additional cue starts another 1.2 s late beyond the shared offset, but its end carries only the shared 0.8-second delay.

The creator first applies the global 0.8-second shift and exports again without editing the anomalous cue separately. Playback of the new render then re-observes all four checkpoints. The three representative cues now have zero start and end error. The anomalous cue now starts at 03:07.2 and ends at 03:09.4, leaving its start +1.2 s late and its end error at zero. This new evidence supports changing only that cue's start to 03:06.0 while retaining the observed end at 03:09.4. The creator then checks the preceding cue's end, this cue's start and end, and the following cue's start for an unintended overlap, a truncated display interval, or a long gap.

After the repair, the creator exports again, replays every start and end checkpoint, and then watches the video in sequence. These numbers illustrate this fictional record only; use boundaries observed in your own final media for a real project. If both start and end errors instead grow through the runtime from +0.2 to +1.1 and then +2.0, do not copy the 0.8-second global shift. Follow the drift path: verify the versions and rebuild the timing.

This example chooses speech boundaries as its target. For a project with accepted reading allowance or deliberate speech across cuts, first use its presentation plan to identify unwanted errors. Zero and 0.8 seconds here do not establish an industry tolerance or mandatory standard.

Replay the Final Render; Recheck Narration Changes and Timing-Affecting Edits

A timing table records the intended boundaries; the rendered media shows the actual cue start and end and how those boundaries feel against the spoken start and end. Finish with this timing-only checklist:

  • This pass uses the video intended for release, the narration actually adopted, and the complete SRT, with no old version mixed in.
  • Speech start and end plus cue display start and end near the beginning, middle, and end have been replayed in the new render.
  • Every local anomaly and the boundaries of its neighboring cues have been replayed, with no unintended overlap, truncated display interval, or long gap.
  • Checkpoints were added around transitions, sections, local speed changes, or recuts that change subtitle-relevant timing or order.
  • The video was played in sequence from start to finish to confirm that cues appear and end in the intended order, without lingering into the next line or disappearing early.
  • The file being published is the same version that was reviewed.

Recheck the relevant boundaries in a new render whenever narration is rerecorded, rewritten, repaced, or given different pauses. If that change alters total duration, recheck later checkpoints too. A video edit invalidates the corresponding timing evidence only when it changes subtitle-relevant timing, order, duration, or timeline position. Color grading, cover changes, and other edits that leave the timeline unchanged do not require a new subtitle-timing pass by themselves.

FAQ

Why does my SRT start in sync and drift later?

That pattern points to growing drift rather than a simple starting offset. Compare speech and cue start and end near the beginning, middle, and end, then verify that the SRT, narration, and video come from the same final version. If either boundary error grows through the runtime, rebuild that timing against the final narration; one global shift can correct only one position.

Can AI tell whether subtitles are synchronized from the SRT text alone?

No. Text alone does not reveal when speech starts and ends in the final media or when a rendered cue appears and disappears, so it cannot prove synchronization. AI can organize manually verified boundary times, calculate errors where both values exist, and suggest a review scope, but unknowns must remain unknown.

Do I need to recheck subtitles after changing one narration line?

Recheck the changed line’s speech and cue start and end in the new render. If the replacement changes duration, pauses, or later cue positions, also remeasure before the edit, at the edit, and near the end. A color-only video edit that leaves subtitle-relevant timing unchanged does not require a new timing pass. Observations from the old narration do not prove the new narration is synchronized.

Do it with the skill

When using /video-recap (or $video-recap in Codex) to organize a subtitle-sync problem, provide the release video version, the narration version actually adopted, and manually verified speech and cue start and end times near the beginning, middle, and end. Ask it to classify only the observed evidence as global offset, growing drift, or local boundary error and to preserve unknowns. Do not ask it to invent timestamps or claim synchronization from text alone.

Read the method: Output clock, cue boundaries, and track-validation scope

About Video Recap SkillsThe video-recap skill on GitHub

All guides