Let a coarse transcript point you to the neighborhood
Automatic transcript ranges may come from fixed chunks, so a sentence can cross a chunk boundary. Text in a range makes that neighborhood worth reviewing; it does not mean the first word begins at the range start or the last word ends at its finish. An empty result also does not prove the absence of quiet dialogue, action sound, or a recognition failure.
Before editing, confirm that the transcript describes the current source file. If the video, selected audio, or name glossary changed, an old result is only a clue from an older version. Once you find the target phrase, listen from before the range until the sentence, pause, and relevant response are clear, then record the actual in and out points you plan to use.
Copyable coarse-transcript-to-edit worksheet
| Item | What to record |
|---|---|
| Current source | <File and version; confirm the transcript belongs to it> |
| Coarse ASR range | <A search starting point only> |
| Rough text | <Raw transcript; flag names, negation, and fragments> |
| Listening range | <How far before and after you listen> |
| Verified phrase | <Confirmed words; mark anything unclear> |
| Source in and out | <Chosen around complete sound, action, and response> |
| Must preserve | <Condition, negation, speaker response, or action result> |
| Current output position | <Read after editing; never guessed from source time> |
| Downstream binding | <Which current clock captions, narration, or effects use> |
Use one row per target phrase. Complete source and listening checks before setting cuts. Leave the output position blank until the current edit exists.
Original example: three complete instructions across four rough windows
The following repair-bench video and paper timestamps are invented for this lesson. No real ASR run, listening session, or edit occurred. In the video, repairer Cen Ning is handing two similar storage boxes to colleague Jiang Xu. The blue-tag box contains logged adapters; the green-tag box still awaits an inventory.
Suppose the coarse transcript says:
| Coarse ASR range | Rough text |
|---|---|
| 0.00–6.00 | “Keep the blue-tag box at the desk” |
| 6.00–12.00 | “until Friday” |
| 12.00–18.00 | Empty text |
| 18.00–24.00 | “collect it after five” |
The table does not show where the first phrase ends or whether the empty range is quiet. For this paper demonstration, listening across 0.00–12.00 reveals a complete line at 4.80–7.90: “Keep the blue-tag box at the desk. Don’t send it out before Friday.” After Cen finishes, Jiang moves his hand away from the blue box. Listening across 12.00–18.00 reveals Cen speaking quietly at 14.20–15.10: “The green one hasn’t been inventoried.” Within 18.00–24.00, Jiang says at 19.10–21.20, “Then I’ll collect the blue one after five.”
These boundaries and lines are entirely invented teaching material. A real project must listen to its own current file instead of copying these numbers.
Choose source cuts around complete meaning
If this example keeps only 0.00–6.00 for the first line, it loses the condition about Friday and clips the sentence. If 12.00–18.00 is discarded as silence, the distinction that the green box still needs an inventory disappears. A more complete paper selection is:
| Clip | Source range and reason |
|---|---|
| A | 4.70–8.40; preserve the full condition for the blue box and Jiang’s current response of moving his hand away |
| B | 14.10–15.40; preserve the complete green-box clarification without merging it into the blue-box rule |
| C | 19.00–21.50; preserve Jiang’s collection response and its reference to the blue box |
These boundaries remain author choices for this example, not a fixed padding formula. In real material, adjacent speakers, a decisive action after the line, or an existing picture cut can change the required range. ASR gets the editor close; listening to and watching the current source determines the final boundary.
Record output time again after the edit
Suppose this example arranges A, B, and C in that order after other approved footage. Source time 4.70–8.40 does not automatically become output time 4.70–8.40. Deletions, reorderings, retiming, and frame boundaries change output positions. After the current edit exists, read the three locations from its actual timeline, then bind captions, narration, and effects to that output clock.
Changing A’s in or out point also moves the output positions of B and C. Applying one offset to every caption cannot represent a local duration change. Record old-edit output time, source time, and new-edit output time separately instead of combining them in one field called time.
A creator workflow from rough transcript to current cut
- Lock the current source file, selected audio track, and matching transcript.
- Use keywords to find a coarse range and inspect its neighboring ranges.
- Listen for actual words, speaker, and phrase boundaries; listen to empty-text ranges too.
- Check picture, action, and response while choosing complete source in and out points.
- Record conditions, negation, references, and results that must survive.
- After producing the current edit, read its real output positions.
- Bind captions, narration, and effects to the current output clock, then play the full sequence and listen without picture.
If listening shows that the source itself is incomplete, leave it pending for replacement, rerecording, or a rewrite. Complete captions cannot create missing audio, and a fade cannot restore a word absent from the source.
Five common misuses and direct repairs
| Misuse | Direct repair |
|---|---|
| Treating chunk start as the first word | Listen earlier and find the actual phrase start and previous boundary |
| Treating chunk end as the sentence end | Listen later and preserve conditions, negation, and final words |
| Treating empty text as silence | Play the range and separately record speech, action sound, and genuine quiet |
| Reusing old output time after an edit | Read positions from the current cut and update downstream bindings |
| Listening to words without viewing responses | Check speaker, referenced object, completed action, and listener response with picture |
Repair boundaries that change meaning before removing wait time with no function. A paper plan is not evidence of listening or export; only playback and inspection of the current file support those results.
FAQ
Can I place narration wherever an ASR window has no text?
Listen to the current track first. Empty text can reflect quiet speech, action sound, recognition failure, or genuine quiet. Decide whether source sound, action, silence, or narration owns the interval only after checking it.
Do it with the skill
Use $video-recap on my current video, run only the understanding and analysis stage, and stop after producing the coarse transcript, understanding indexes, and creative brief. List the rough ranges containing target dialogue and state that they are not exact phrase boundaries; do not treat empty text as proof of silence. Listen to the current source before setting cuts, then record output time separately after editing.
Read the method: Public video-recap entry point and analysis-stage pause · Video-understanding workflow and coarse-transcript limits · ASR coarse ranges, empty text, and timeline data · Implementation of coarse-timing precision and current-file binding · Source, dialogue, and timeline checks before scripting