Separate a video analysis into four layers
Layer one records direct observation: who enters the image, what a hand touches, where an object moves, and which words are actually visible on screen. Layer two records dialogue that was heard, while preserving an unidentified speaker when necessary. Layer three adds interpretation, such as how an action changes access to an object; its strength cannot exceed picture and sound evidence. Layer four lists unknowns, including motive, off-screen agreements, verified identity, later results, and authorial intent.
A frame supports the visible state at that instant, not an entire action by itself. Automatic captions can help locate speech, but they do not establish exact wording, names, or speaker identity. Background notes can supply relationships, but they cannot declare what happens in the current shot.
Copyable two-column video-analysis worksheet
| Item | What to record |
|---|---|
| Current question | <The person, action, or change this excerpt must establish> |
| Visible facts | <Positions, object states, action boundaries, and visible text> |
| Heard dialogue | <Exact line or pending fragment; state whether the speaker is known> |
| Supported change | <Actual change in knowledge, goal, relationship, power, emotion, or risk> |
| Limited interpretation | <The picture and dialogue that jointly support it> |
| Cannot establish | <Motive, identity, permission, result, or off-screen event> |
| Review range | <Premise, action, response, and local outcome> |
| Final wording | <Two or three specific, traceable sentences that preserve unknowns> |
Write the question first, then gather evidence from the complete excerpt. If a limited interpretation has no matching image or line, move it to cannot establish or remove it.
Original footage sketch: the tool case leaves the table
The following paper footage was invented for this lesson. It was not filmed, analyzed by a model, or play-tested. A community repair event takes place in an outdoor corridor. The check-in table stands at the edge of an awning, with a covered shelf against the wall behind it. At the start, volunteer Mei Qing puts a red case labeled SCREENING-ROOM ADAPTERS on the table. The rain becomes louder, and drops begin hitting one corner of the table.
Lin Tao enters the image, points to the case, and says, “Water is coming in here. I’ll move it under the awning.” He lifts it with both hands, walks continuously along the same corridor, and places it on an empty chair under cover. The footage does not show him opening it or establish whether he is responsible for storage.
Later, Mei returns to the check-in table and asks, “Where is the red case?” Another volunteer looks toward the awning. Lin answers from off screen, “Under the awning. I didn’t open it.” Mei walks to the chair and sees the clasps closed. She carries the case to the covered shelf behind the check-in table and says, “Call me before moving it next time.” Lin replies, “Okay.” The excerpt ends with the case on the dry shelf.
The paper footage supports that the case was moved, drops reached a corner of the table before the move, its clasps were closed when Mei inspected them, and the case was then placed on the covered shelf. It does not verify the inventory inside, Lin’s authority, whether nobody touched it off camera, or his private motive.
Repair an overreaching AI summary line by line
An overreaching version might say: “During the confusion, Lin stole the tool case. When discovered, he lied about protecting it from the rain, then returned the stolen property and confessed when confronted with the evidence.” This changes a move into theft, labels his explanation a lie, treats closed clasps as a complete inventory, and expands “Okay” into a confession.
After using the worksheet:
As the rain grew louder and drops reached the table corner, Lin said water was coming in and moved the red case under the awning in the same corridor. When Mei searched for it, he gave its location from off screen. Mei found the clasps closed, placed the case on the covered shelf behind the check-in table, and asked him to notify her before moving it in the future. Lin agreed to that request. The excerpt does not check the contents or establish Lin’s storage authority or a stronger motive.
The repair keeps a readable conflict: Mei needs to know where shared equipment goes, while Lin moves it before informing her. It also preserves the actual result: the case remains on a dry shelf and they make a current agreement about future notice.
Do not let one frame narrate a complete action
A single frame of Lin holding the case can turn possession at one instant into a claim of ownership or theft. Review the condition before the action, the continuous action, when another person receives information, the response, and the local outcome. If the footage cuts, record what appears on each side instead of joining different places or times into one continuous fact.
Use the same boundary for emotion. A frown, pause, or averted gaze is observable. Guilt, jealousy, and prior intent are interpretations requiring dialogue, a behavioral sequence, or reliable context. When evidence is thin, “Mei pauses and looks toward the awning” is more accurate than supplying her private thoughts.
Run five reverse checks before delivery
- Circle every word that assigns motive, identity, or outcome; point each one to picture, dialogue, or reliable context.
- Temporarily remove deep interpretation and read only direct observations; confirm they still describe the excerpt accurately.
- Return to the complete excerpt and check for a premise, response, or local result that changes the meaning.
- Confirm visible text and automatic captions against the actual material; leave unclear words pending.
- Confirm background research supplies context without proving current action.
These checks improve traceability. They do not turn a paper analysis into a completed review of real footage. A live project still requires watching the current source and final output.
FAQ
Can a detected facial expression establish a character’s motive?
An expression can be recorded as a visible response. Motive still needs dialogue, a sequence of behavior, or reliable context. Keep the expression, interpretation, and unknowns separate instead of turning one frown into a claim about crime, deception, or a lasting relationship.
Do it with the skill
Use $video-recap on my video, run only the understanding and analysis stage, and stop after producing the understanding indexes and creative brief. For each excerpt, separate visible facts, heard dialogue, evidence-backed interpretation, and unknowns; attach a source range to each interpretation. Preserve uncertainty around motive, identity, relationships, and outcomes, and do not treat an empty transcript window as proof of silence.
Read the method: Public video-recap entry point and analysis-stage pause · Fact, inference, and uncertainty boundaries in video understanding · Frame observation and deeper-analysis prompt template · Data boundaries for frame facts and coarse transcription