First, what a reference breakdown is for
video-reference is an on-demand breakdown skill outside the default production path. It splits a finished video into two layers: how this video was made (facts) and what would still work with different footage (methods). Facts stay on your machine; the exported production_reference.json holds only methods and measured numbers.
It calls neither MiMo nor any other LLM, scores nothing and does not judge how far your new edit is from the reference. The next time you write, video-script reads the file only if it is in work_dir, with this priority: user instructions, then evidence from your own footage, then the reference. The numbers are reference values, not quotas. When they conflict with sound ownership, complete lines or performances in your footage, the footage wins.
The editing rhythm article weighs individual cuts. This one starts with a set of checkable numbers and then decides, method by method, what to use.
Five steps: measure, review, label, write methods, export
The input is the finished file, ideally with video-understanding output produced for it (5-second ASR windows are recommended). Without that output the breakdown still runs, but narration speed stays empty. Video understanding is a separate skill whose own instructions say whether it goes online; video-reference itself only decodes locally with ffmpeg.
- measure: one pass over the whole video writes cuts, the shot-length distribution and loudness; an unchanged file reuses the cache.
- frames review: a fixed threshold is wrong in both directions on real videos, so every run needs a look at the frame sheets for flagged windows (
--review) and the five longest shots (--longest 5). Record missed and false cuts incut_fixes; if you looked and nothing needs changing, write{}. - Label: who owns the sound (narration, original dialogue, action sound, ambience, music, silence) and each section’s story function (hook, setup, turn, escalation, payoff). Place sound boundaries where the sound actually starts and stops, not at the subtitles.
- Write facts and methods: each fact has a time or measurement anchor. Each method gives a rule, when to use and avoid it, which planning file it serves and its evidence; numbers only point at measurement paths and are never typed by hand.
- check → export: export only with zero errors. Narrative structure, pacing, shots and editing, narration and subtitles, and sound-picture handoff each need at least one method or a stated reason for skipping.
What gets measured, and how precise it is
| Value | Where it comes from |
|---|---|
| Cuts and shot lengths | ffmpeg scene scores: 10 or above always counts; 4 or above counts only as an isolated peak within 0.3 seconds. Gives median shot length, p10/p90, share under 1 second and over 8 seconds, cuts per minute and a 10-second curve. reviewed after your frame check, otherwise measured |
| Loudness | Integrated loudness, loudness range (LRA), true peak and 1-second short-term loudness; measured |
| Each sound owner | Seconds, share, median block length, cuts per minute inside it, mean loudness; narration can also be split by job (context, causal link, foreshadowing, interpretation, transition) |
| Each story section | Seconds, cuts per minute and narration share, with positions as fractions of the running time |
| Narration speed | Characters per second from ASR windows at least 80% covered by narration; coarse-window precision, not word alignment |
| Handoffs on cuts | Share of sound-ownership changes within ±0.25 seconds of a picture cut |
| First original dialogue | Where original dialogue (including the source’s own voice-over) first appears, exported only as a fraction |
| Subtitle form | Burned in or not, maximum lines, how original lines are marked |
The last six come from your labels (exported pacing values are marked labeled) and are only as accurate as those labels.
What it cannot measure, and what you cannot take
What it cannot measure or judge: slow dissolves go undetected, and a hard cut into darkness after motion or two similar-scoring cuts within 0.3 seconds are not counted automatically; they only enter the review windows. The checks do not judge meaning, so you must make each method’s direction agree with the derived values. A number shows only that the reference did something, not that it worked or that it suits your material.
What you cannot take:
- Before export the check blocks source names, sentences sharing 8 consecutive Chinese characters with dialogue or research, absolute timecodes and paths. It catches only literal leaks, not a paraphrased plot.
- The breakdown files and frame sheets in
U/contain source facts and images. Keep them local, out of new projects and out of commits. - The reference’s footage, narration text and section-by-section layout do not go into your video, and you should not rewrite its narration line by line. One of the skill’s own bad examples is “follow this video’s structure”: it has no conditions and cannot transfer.
- You need your own local copy of the reference. Whether you may obtain, analyze and use it is yours to confirm.
This matches the boundary in analyzing a reference drama’s audiovisual choices: transfer the approach, not the content.
Original paper example: a reference profile against a ferry recap draft
Everything below is invented for this article. Reference R is an imaginary 6-minute recap of a mystery short and matches no real channel, work or video. The numbers are illustrative only; nothing was actually run, edited or listened to.
Two methods exported from R (illustrative):
- m1 | narrative structure, for the story plan: one opening narration line sets up the situation, then the first original-dialogue conflict takes over. Use when the conflict line can be followed on its own; avoid when it needs a lot of backstory.
- m2 | sound-picture handoff, for the visual-audio board: longer shots under narration, the source’s own cutting rhythm in original-dialogue sections. Avoid in sections driven by action sound.
My project is a licensed 8-minute original short. Before the last ferry leaves, the ticket clerk A-Cheng notices that a passenger’s ticket is printed with tomorrow’s date; the argument plays as one long two-shot. I estimated my draft’s numbers by hand from my own shot list:
| Item | Reference R (illustrative) → my draft (by hand) |
|---|---|
| First original dialogue | 0.04 → about 0.19 (about 95 opening seconds of narration) |
| Narration share | 0.55 → about 0.72 |
| Cuts per minute, narration | 12 → about 15 (many rope and lamp inserts) |
| Cuts per minute, original dialogue | 24 → about 8 (the source’s long take) |
| Narration speed | 4.1 characters/s (coarse windows) → about 4.6 (script characters ÷ planned time) |
I made this comparison by hand; it is not skill output. Decisions, one method at a time:
- Adapt m1: cut the opening narration to one line, “Ten minutes before the last ferry, A-Cheng is alone at the window,” go to the passenger handing over the ticket, and bring her line “This ticket is for tomorrow” forward to about 0.07. I stop short of 0.04 because viewers must see the handover before her line makes sense.
- Use half of m2: drop the repeated inserts under narration, bringing that rate to about 11. Skip the reference’s 24 for original dialogue: A-Cheng’s pause before she stamps the ticket survives only if I do not cut. The footage decides.
Narration share falls to about 0.62, and I stop there instead of cutting more to approach 0.55. Narration speed waits until the voice-over exists and I have heard it at normal speed. Both decisions go into the optional reference_methods field in recap_story_plan.json as adapt, each with its reason. This changed two plan decisions; it does not prove the new draft plays better, and none of R’s section order, narration or images entered this draft.
Five common misuses
- Matching every reference value: first ask whether your footage meets the method’s conditions, then look at the number.
- Reading only whole-video cuts per minute: the overall average blurs the difference between original-dialogue and narration sections. Read the values split by sound owner and by section.
- Using shot counts without review: a hard cut into darkness may score only 5–7, while one fast-moving shot can produce a 7–8 peak every 0.2 seconds. If a new measurement changes the review windows, look at the new windows again.
- Setting sound boundaries from subtitles: narration end points taken from subtitle frames were measured 0.24–0.42 seconds late, while handoff alignment tolerates only ±0.25 seconds, so misplaced boundaries make that share reflect labeling habits.
- Writing methods as plot summaries: entries with names, retellings of events or specific time points cannot transfer. Rewrite them as “under these conditions, do this.”
FAQ
Can I apply the reference’s section structure directly to my draft?
Not wholesale. In free creation (CREATE) mode the writer may treat the reference structure as one candidate hypothesis and compare it with hypotheses grown from your own footage; in directed or revision modes (DIRECTED / REVISION) it is not applied by default. Even when chosen, you borrow only section functions and which sound leads, never the content.
Can video-reference compare my finished video with the reference and score the gap?
No. It scores nothing, calls no model and makes no judgment about the gap between a new video and the reference. No script reads the export, and its presence does not change whether any stage passes. The comparison table in this article was worked out by hand from the writer’s own shot list.
Can I break down a reference without ASR or understanding output?
Yes for cuts, shot lengths and loudness, and you can still label sound ownership and sections. Narration speed stays empty, the name-leak scan only uses the entities you write in facts, and the check warns you. To view the picture, use a contact sheet sampled every 2 seconds across the whole video.
Do it with the skill
Use $video-reference to break down the reference recap I provide locally (plus its video-understanding output folder, if any). Run measure, view the frame sheets for flagged windows and the longest shots, and record cut_fixes; label sound ownership and story sections at the actual sound boundaries, write source facts separately from transferable methods, and point targets only at measurement paths. Export to my next project’s work_dir only after check reports zero errors. Do not score either video or judge the gap between my edit and the reference; keep source names, dialogue and timecodes out of the export, and keep the breakdown and frame sheets local.
More of the method: How the writing stage weighs a reference and records reference_methods