Let a coarse transcript point you to the neighborhood

Automatic transcript ranges may come from fixed chunks, so a sentence can cross a chunk boundary. Text in a range makes that neighborhood worth reviewing; it does not mean the first word begins at the range start or the last word ends at its finish. An empty result also does not prove the absence of quiet dialogue, action sound, or a recognition failure.

Before editing, confirm that the transcript describes the current source file. If the video, selected audio, or name glossary changed, an old result is only a clue from an older version. Once you find the target phrase, listen from before the range until the sentence, pause, and relevant response are clear, then record the actual in and out points you plan to use.

Copyable coarse-transcript-to-edit worksheet

ItemWhat to record
Current source<File and version; confirm the transcript belongs to it>
Coarse ASR range<A search starting point only>
Rough text<Raw transcript; flag names, negation, and fragments>
Listening range<How far before and after you listen>
Verified phrase<Confirmed words; mark anything unclear>
Source in and out<Chosen around complete sound, action, and response>
Must preserve<Condition, negation, speaker response, or action result>
Current output position<Read after editing; never guessed from source time>
Downstream binding<Which current clock captions, narration, or effects use>

Use one row per target phrase. Complete source and listening checks before setting cuts. Leave the output position blank until the current edit exists.

Original example: three complete instructions across four rough windows

The following repair-bench video and paper timestamps are invented for this lesson. No real ASR run, listening session, or edit occurred. In the video, repairer Cen Ning is handing two similar storage boxes to colleague Jiang Xu. The blue-tag box contains logged adapters; the green-tag box still awaits an inventory.

Suppose the coarse transcript says:

Coarse ASR rangeRough text
0.00–6.00“Keep the blue-tag box at the desk”
6.00–12.00“until Friday”
12.00–18.00Empty text
18.00–24.00“collect it after five”

The table does not show where the first phrase ends or whether the empty range is quiet. For this paper demonstration, listening across 0.00–12.00 reveals a complete line at 4.80–7.90: “Keep the blue-tag box at the desk. Don’t send it out before Friday.” After Cen finishes, Jiang moves his hand away from the blue box. Listening across 12.00–18.00 reveals Cen speaking quietly at 14.20–15.10: “The green one hasn’t been inventoried.” Within 18.00–24.00, Jiang says at 19.10–21.20, “Then I’ll collect the blue one after five.”

These boundaries and lines are entirely invented teaching material. A real project must listen to its own current file instead of copying these numbers.

Choose source cuts around complete meaning

If this example keeps only 0.00–6.00 for the first line, it loses the condition about Friday and clips the sentence. If 12.00–18.00 is discarded as silence, the distinction that the green box still needs an inventory disappears. A more complete paper selection is:

ClipSource range and reason
A4.70–8.40; preserve the full condition for the blue box and Jiang’s current response of moving his hand away
B14.10–15.40; preserve the complete green-box clarification without merging it into the blue-box rule
C19.00–21.50; preserve Jiang’s collection response and its reference to the blue box

These boundaries remain author choices for this example, not a fixed padding formula. In real material, adjacent speakers, a decisive action after the line, or an existing picture cut can change the required range. ASR gets the editor close; listening to and watching the current source determines the final boundary.

Record output time again after the edit

Suppose this example arranges A, B, and C in that order after other approved footage. Source time 4.70–8.40 does not automatically become output time 4.70–8.40. Deletions, reorderings, retiming, and frame boundaries change output positions. After the current edit exists, read the three locations from its actual timeline, then bind captions, narration, and effects to that output clock.

Changing A’s in or out point also moves the output positions of B and C. Applying one offset to every caption cannot represent a local duration change. Record old-edit output time, source time, and new-edit output time separately instead of combining them in one field called time.

A creator workflow from rough transcript to current cut

  1. Lock the current source file, selected audio track, and matching transcript.
  2. Use keywords to find a coarse range and inspect its neighboring ranges.
  3. Listen for actual words, speaker, and phrase boundaries; listen to empty-text ranges too.
  4. Check picture, action, and response while choosing complete source in and out points.
  5. Record conditions, negation, references, and results that must survive.
  6. After producing the current edit, read its real output positions.
  7. Bind captions, narration, and effects to the current output clock, then play the full sequence and listen without picture.

If listening shows that the source itself is incomplete, leave it pending for replacement, rerecording, or a rewrite. Complete captions cannot create missing audio, and a fade cannot restore a word absent from the source.

Five common misuses and direct repairs

MisuseDirect repair
Treating chunk start as the first wordListen earlier and find the actual phrase start and previous boundary
Treating chunk end as the sentence endListen later and preserve conditions, negation, and final words
Treating empty text as silencePlay the range and separately record speech, action sound, and genuine quiet
Reusing old output time after an editRead positions from the current cut and update downstream bindings
Listening to words without viewing responsesCheck speaker, referenced object, completed action, and listener response with picture

Repair boundaries that change meaning before removing wait time with no function. A paper plan is not evidence of listening or export; only playback and inspection of the current file support those results.

FAQ

Can I place narration wherever an ASR window has no text?

Listen to the current track first. Empty text can reflect quiet speech, action sound, recognition failure, or genuine quiet. Decide whether source sound, action, silence, or narration owns the interval only after checking it.

Do it with the skill

Use $video-recap on my current video, run only the understanding and analysis stage, and stop after producing the coarse transcript, understanding indexes, and creative brief. List the rough ranges containing target dialogue and state that they are not exact phrase boundaries; do not treat empty text as proof of silence. Listen to the current source before setting cuts, then record output time separately after editing.

Read the method: Public video-recap entry point and analysis-stage pause · Video-understanding workflow and coarse-transcript limits · ASR coarse ranges, empty text, and timeline data · Implementation of coarse-timing precision and current-file binding · Source, dialogue, and timeline checks before scripting

About Video Recap SkillsThe video-recap skill on GitHub

All guides