What the index contains and what each file answers

video-understanding is the analysis stage that runs before writing. Its role is footage observer, a script supervisor rather than a director: it turns the source video into a set of files for the writing stage and does not write narration itself.

FileWhat it answers and how to use it
scenes.jsonCut points and each scene’s start and end, with junk segments filtered; a scene is not a story beat
asr_result.jsonDialogue text in coarse windows; use it only to find roughly where speech is
asr_timing_evidence.jsonWhether the transcript is usable, its precision limits, and text before and after name-list correction
silence_periods.jsonQuiet windows with has_speech; candidate places for narration
vlm_analysis.jsonPer-scene description, depth analysis, and timestamped frame_facts
timeline_fusion.jsonVisuals, dialogue, and quiet windows aligned scene by scene
agent_narration_brief.mdThe writing brief read first before scripting

There is also a source contact sheet, storyboard/source_storyboard.jpg, and a global story index, understanding_index.md, which is produced by default and is also a model-written summary.

Before you run it: rights, dependencies, and background research

  • Rights: process only video you own or have permission to use. Extracted frames and audio are sent to a remote model service, so confirm your rights and read the provider’s privacy and billing terms first.
  • Local tools: run scripts/understand.py with python3; ffmpeg is required. Intermediate files go into --work-dir.
  • Model services: transcription uses mimo-v2.5-asr and visual observation uses mimo-v2.5; both need MIMO_API_KEY. --skip-asr skips only transcription and overwrites any existing transcript with [], so it is no way around a missing key.
  • Background research (optional): the synopsis and character names in work_dir/background_research.json are folded into visual analysis, and --context can add one more hint. Descriptions can then name people, but a name can also land on the wrong person. For how to research, see How to Research Before Writing a Video Recap.
  • Caching: outputs are reused only when they match the current video and relevant settings; add --force to recompute everything.

Check the index in this order before writing narration

  1. Read the status first. The brief prints the transcript evidence status: AVAILABLE_COARSE means a coarse transcript exists; UNAVAILABLE_NO_KEY, FAILED_PROVIDER, and similar values mean this run has no usable dialogue; MISSING_OR_STALE means the evidence is missing or does not match the current files. When transcription or visual observation is thin, you get a material warning; draw fewer conclusions.
  2. Check scene splits on the contact sheet. A missed cut merges two scenes into one description; an extra cut splits one action apart. Make sure nothing filtered out is a shot you need. If the splits are wrong, adjust --scene-threshold and rerun.
  3. **Compare each description with its frame_facts.** Every action, object movement, and name in a description should appear in some frame tag. Tags cover only the sampled frames; what happens between two of them is filled in by the model.
  4. Treat depth analysis as a hypothesis. The prompt template asks the model to read emotion, relationships, and subtext from visible evidence and to say so when information is insufficient. It is still the model’s judgment.
  5. Check words that change the meaning. Listen first to names, negatives, and numbers; for windows the name list altered, compare both versions. Empty text is not proven silence.
  6. Look for contradictions across files. The fused timeline sets each scene’s visuals beside its dialogue. When a later scene’s dialogue contradicts an earlier description, rewatch there first.
  7. Rewatch every beat you will rely on. Record findings in your own check notes, marked as rewatched by hand, and never pass hand edits off as the model’s original output.

Original example: one wrong observation in a three-minute market short

The short film, people, and times below are invented for this article; no model was actually run. The JSON is hand-written for illustration, borrows field names from data-schema, omits some fields, and is not real output from the skill.

Footage: *The Borrowed Scale*, a three-minute short the author filmed at a morning market. The fish stall’s scale has broken, so the owner, Aunt Miao, sends her niece Xiaohe to borrow one from Old Qian’s stall next door. Qian wants it back before noon. Around 2:06, Xiaohe carries the scale into the aisle between the stalls; around 2:38, Qian comes over and asks where it is; at the end, Xiaohe takes the scale from under her own counter and hands it to him.

An excerpt from the fused timeline (scene_id counts from 0, so the brief shows scene 5):

{

"scene_id": 4,

"time_range": [126.0, 152.0],

"visual_description": "Xiaohe carries the scale to Qian and returns it",

"depth_analysis": "She returns it before noon as agreed",

"frame_facts": {

"128.0": ["Xiaohe carries the scale into the aisle"],

"136.0": ["Xiaohe stands between the stalls, no scale in hand"]

},

"narration_slots": [{"start": 141.0, "end": 149.0, "duration": 8.0}]

}

One transcript line from the next scene (asr_result.json):

{"start": 158.0, "end": 168.0, "text": "Where's the scale? You said before noon."}

Written straight from the index, the quiet window at 141–149 s could easily become “The scale went back on time.” Checking turns up the problem in three places:

  • The description goes past the frame tags. The tags say only “carries the scale” and “no scale in hand.” No frame shows Qian taking it; “returns it” is the model’s fill-in.
  • The files contradict each other. In the next scene someone asks where the scale is. The transcript does not mark speakers, so note it as “someone asks about the scale.”
  • The footage settles it. In 126–140 s, Xiaohe sets the scale on the corner of her own counter and turns to serve a customer; Qian is not in frame.

Check note: scene 5 becomes “Xiaohe leaves the scale on the corner of her own counter and does not hand it to Qian” (rewatched 126–140 s); the “as agreed” analysis is dropped; the question near 158 s is Qian’s, quoted only after listening word by word.

The story changes with it: a promise kept becomes a promise to return the scale before noon, forgotten and put right only at the end. Whether that quiet window gets narration is the author’s call after watching the footage. If it does, it says only what the picture proves: the scale is still on their own counter.

What the index cannot tell you

  • Dialogue boundaries and speakers. Transcript start/end values are coarse windows formed from fixed chunks, not word alignment, and they do not say who is speaking; see How to Edit Video With Inaccurate ASR Timestamps.
  • Action between frames. Hand-offs, hiding an object, and who moved first need continuous viewing.
  • Motive, off-screen events, and what happens afterward. Whatever the footage did not capture is unknown; see How to Stop AI Video Analysis From Inventing the Story.
  • Whether a silence should stay. has_speech set to false only means the stretch is quiet and does not overlap detected transcript speech. Whether the silence is itself the performance is your call.
  • Whether names are right. Names may come from background research or the name-list correction, and people who look or dress alike get mislabeled.
  • Story direction. The style in the brief is just the --style value, documentary by default. The viewer question and point of view belong to the writing stage; see How to Write a Recap Script.
  • Whether you may use the footage. The index does not judge copyright or permission.

Common mistakes and fixes

MistakeFix
Reading only visual_descriptionCompare it with frame_facts; rewatch any action it adds
Writing depth_analysis as factRewrite it as visible action, or find supporting dialogue and continuous behavior
Filling every quiet window with narrationFirst decide whether the silence is doing work
Keeping an old index after the footage changedRerun on the current file; preserve file times, for example with cp -p, when moving a work directory
Running --brief-only to fix a visual observationIt rebuilds the brief from existing outputs and does not look at the frames again

FAQ

Can I build the index without a MiMo API key?

A full run needs MIMO_API_KEY, because both transcription and visual observation call MiMo models. Without a key, a rerun only reuses cached transcripts, visual analysis, and story index, and any summary step that would call a model is recorded as skipped. Do not use --skip-asr as a workaround: it replaces an existing transcript with an empty list, and visual observation still needs the key.

The index names people. Can I use those names in the narration?

Check them first. Names may come from background_research.json or --context folded into the visual analysis, or from name-list corrections in the transcript, and people who look or dress alike are easily mislabeled. Confirm each name from on-screen text, how characters address each other, or a cast list from the creator. If you cannot confirm it, refer to the person by role or action.

If the footage was re-exported or re-edited, is the old index still usable?

Treat it as new footage. Each stage is reused only when its outputs and provenance records match the current video and relevant settings. The transcript evidence records the size and modification time of the source video, audio, and transcript, and the brief shows MISSING_OR_STALE when they no longer match. Times and descriptions in an old index are only leads about the old version.

Do it with the skill

Use $video-recap to analyze this local video that I have the right to use, running only through its video-understanding stage and stopping at the pause after the writing brief: produce the scenes, transcript, quiet windows, per-scene observations, fused timeline and writing brief, then stop without writing narration. Report the transcript evidence status and any material warnings first, then list, scene by scene, the actions and names in descriptions that no frame fact supports, so I can check them against the footage.

More of the method: Prompt for scene descriptions, frame tags, and depth analysis · How background research feeds visual analysis

Install and use Video Recap Skills

All guides