The Image-to-Video Prompt Formula: Start with an Information-Deletion Table
An image-to-video request already includes a real reference image. Decide item by item whether the image can carry the information or the motion prompt must still carry it. This prevents the prompt from becoming a second static-image description.
| Item | Details |
|---|---|
| Face, hairstyle, complete outfit | Image: Yes; Motion condition: A local feature participates in the action or must remain fixed; Example: The left hand grips the right cuff; the sleeve remains rolled below the elbow at the end |
| Set decoration and baseline lighting | Image: Yes; Motion condition: The environment changes or constrains the path; Example: The door opens inward, so the character must turn sideways between the table corner and the door |
| Static composition and starting positions | Image: Yes; Motion condition: Camera motion or a character’s relocation changes the relationship; Example: The character moves from frame left through the doorway; the empty table remains visible on the right at the end |
| Hands, held objects, and contact | Image: No; Motion condition: These directly determine whether the action is executable; Example: The left hand supports the saucer; the right hand holds the pot handle; after pouring, each hand keeps its own object |
| Gaze and object of attention | Image: No; Motion condition: They trigger performance or define the end state; Example: Look at the signature first, then raise the eyes toward the person offering the paper |
| Occlusion and off-screen entry path | Image: No; Motion condition: An object disappears and then reappears; Example: The folder moves from behind the cabinet door; specify the order in which it becomes visible |
| Visible text and fixed markings | Image: Depends on the project; Motion condition: The text-bearing surface moves, becomes occluded, or must stay fixed; Example: The label remains centered on the lid and turns with the lid’s surface |
Deletion test: remove one static description, then ask, “Can the reference image answer this in only one way while preserving the action’s result?” If yes, delete it. If deletion makes the active hand, contact object, occlusion path, or endpoint ambiguous, retain the shortest motion fact that resolves the ambiguity.
Compare Start and End States: Write Only What Changes
Put the reference-frame start state and the storyboard endpoint into one difference table. Action sentences explain how every difference occurs. Mention an unchanged field only when it is likely to drift.
| Item | Details |
|---|---|
| Posture / center of gravity | Start: Seated, weight on the chair; End: Standing upright, weight on both feet; Change: Press the right hand on the table, lean forward, then stand upright |
| Gaze | Start: Looking at the photo on the table; End: Looking at the visitor beyond the door; Change: The sound at the door triggers an upward glance, then the head turns toward the doorway |
| Hands / held object | Start: Right hand presses the photo; left hand rests on the table; End: Photo remains on the table; right hand is empty; Change: Release the photo before bracing on the table; do not lift the photo |
| Relative position | Start: In front of the chair and behind the table; End: To the right of the chair and still behind the table; Change: After standing, move half a step toward frame right |
| Visible state | Start: Door closed; visitor off-screen; End: Door open about thirty degrees; visitor remains outside; Change: Door opens first; only the visitor’s shoulder outline enters view |
| Preserve | Start: Water glass on the table; orientation of the photo; End: Both remain unchanged; Change: The glass does not move; the photo remains face up |
Reusable difference card:
Start: posture___; gaze___; left hand___; right hand / held object___; relative position___; visible state___.
End: posture___; gaze___; left hand___; right hand / held object___; relative position___; visible state___.
Preserve: ___.
Change chain: Because of ___, first ___, then ___, and finally ___.
The start and end states come from the existing reference frame and storyboard boundaries. If the endpoint is missing, complete the storyboard first. A prompt should not invent an ending that belongs to story development.
Eight Image-to-Video Motion Problems, with Bad and Revised Prompts
The following examples were designed for this page and have not been sent to a generation service. Each begins with a fact visible in the reference frame, then replaces a static restatement with a prompt that describes only the change.
1. Changing which hand holds an object
Reference frame: The character holds a key in the right hand and a paper bag in the left. Endpoint: The key is in the left hand, and the right hand pushes the door.
- Bad prompt: A young man in a gray coat stands before a wooden door holding a key and paper bag, ready to open the door.
- Revised prompt: The character first sets the paper bag from his left hand beside his feet, then transfers the key from his right hand into his free left hand. Only after the left hand closes around the key does the right hand press the handle and push the door inward about thirty degrees. At the end, the left hand holds the key at waist height, the right hand remains on the door, and the paper bag stays where it was placed.
2. Reappearing after occlusion
Reference frame: A folder protrudes from the left side of a half-open cabinet door. Endpoint: The entire folder lies on the desk.
- Bad prompt: An office contains a cabinet, a desk, and a blue folder. The character takes out the file.
- Revised prompt: The character’s right hand grips the visible top edge of the folder and slides it toward frame right behind the cabinet door. The folder becomes half hidden by the door, emerges completely past its edge, and is placed flat in the center of the desk. The cabinet door keeps the same angle throughout, and the right hand releases the folder at the end.
3. Passing an object between two people
Reference frame: A holds a shallow tray containing three glass balls with both hands. B’s open hands wait on either side of the tray. Endpoint: B securely holds the tray, and all three balls remain in their original grooves.
- Bad prompt: Two people carefully pass a beautiful tray of glass balls with smooth, natural movement.
- Revised prompt: A keeps the tray level at chest height. B first supports the near underside with the left hand, then grips the far rim with the right. Once both of B’s contact points are stable, A releases the right hand and slides the left hand horizontally out from beneath the tray. At the end, the tray retains its height and angle, both of B’s hands bear its weight, and the three balls remain in their grooves.
4. Moving from seated to standing
Reference frame: The character sits behind a table with the right hand pressing a photo. Endpoint: The character stands to the right of the chair; the photo remains on the table.
- Bad prompt: The character sees the doorway and stands up in shock.
- Revised prompt: After one knock sounds beyond the door, the character first raises their eyes, releases the photo with the right hand, and braces that hand on the table edge. They lean forward, stand upright, then move half a step toward frame right to clear the chair back. At the end, the character stands to the right of the chair and looks toward the door; the photo remains face up.
5. Performance triggered by sound
Reference frame: The character looks down while fastening a cuff link. Endpoint: The character looks directly toward an off-screen voice; the cuff link remains in place.
- Bad prompt: Hearing the name, the character looks surprised, then suspicious, with a complicated expression.
- Revised prompt: After an off-screen voice says the full name, the character’s fingers stop on the cuff link and the thumb presses against its face. The character then raises their eyes toward the source of the voice, inhales, and remains silent. At the end, the cuff link has not moved, and the gaze stays fixed toward the sound.
6. Parallel action in the environment
Reference frame: The character’s right palm rests at the center of a fogged car window. A strip of exterior light has not yet entered the frame. Endpoint: A transparent arc crosses the glass, and the light strip passes the arc’s right end.
- Bad prompt: The character wipes fog from the car window while lights pass outside, quiet and atmospheric.
- Revised prompt: As the exterior light strip enters from frame left, the character’s right palm wipes a semicircular arc upward and to the right. The palm’s edge pushes condensation beyond the arc while the light crosses the remaining droplets from left to right and becomes clear inside the wiped area. At the end, the palm stops at the arc’s right end, the transparent arc is continuous, and the light continues out of frame after passing the hand.
7. One motivated camera move
Reference frame: A paper kite lies flat on a roof. The character holds the spool in the right hand, and the string is slack. Endpoint: The kite has risen above the eaves; the character and spool remain in the lower frame.
- Bad prompt: The character launches the kite, and the camera follows it cinematically into the sky.
- Revised prompt: A gust first lifts the kite’s leading edge, and the string pulls taut. The character raises the spool to shoulder height with the right hand as the kite lifts toward the upper right. When the kite passes above the character’s head, the camera makes one upward tilt, moving from the spool to the height relationship between kite and eaves. At the endpoint, the character remains in the lower left, the kite is above the eaves, and a taut string connects them.
8. Ending where the next shot can begin
Reference frame: The character’s left hand holds a doorknob, with the body turned sideways to the door. Start of the next shot: The character is inside the threshold, still holding the knob and looking back at an empty table.
- Bad prompt: The character pushes the door open, enters, looks around, and then keeps walking.
- Revised prompt: After glass clinks beyond the door, the character pushes it open about thirty degrees, leans the head past its edge to look, then steps half a pace across the threshold. The character stops and looks back at the empty table. At the endpoint, the character stands just inside the threshold, the left hand still holds the knob, and the body position matches the start of the next shot.
Local Deformation and Occlusion: Define the Target Area and What Must Stay Fixed
When several objects sit close together in the reference image, instructions such as “make it disappear” or “fold the ground upward” leave the generation process to decide the affected area. Fill these four fields first:
| Item | Details |
|---|---|
| Trigger | Check: Which action or sound begins the change?; Example: The central red line lights from its near end |
| Target area | Check: Which surface changes, along what path?; Example: Only the paper inside the red line folds from near to far |
| End state | Check: What remains, becomes visible, or occupies the area?; Example: A continuous narrow ridge forms down the center |
| Preserve | Check: Which nearby objects keep their position and count?; Example: The four corner weights, paper outside the red line, and table edge remain still |
Completed prompt:
Begin from the reference frame with the paper diagram lying flat and metal weights holding all four corners. After the central red line lights from the near end, only the paper inside that line folds upward along it. The fold advances from near to far and ends as one continuous narrow ridge. The four metal corner weights, the paper outside the red line, and the table edge retain their positions and number. The camera remains fixed overhead, keeping the entire changing boundary visible.
When the occlusion relationship changes, also state what blocks the object first, from which side it reappears, and how far it emerges. This is more direct than stacking generic negative prompts.
Clean Up Before Handoff: Turn Production Notes into Visible Facts
Before handing the prompt to the generation stage, remove filenames, version numbers, internal rule numbers, reasons for rerunning, storyboard-cell numbers, and failure logs. A production note belongs in the prompt only after it has been converted into a visible requirement for the current shot.
| Production note | Visible fact |
|---|---|
| “Use the final outfit” | Clothing matches the reference frame |
| “The previous version had a hand error” | Both hands remain fully visible; the right hand holds only the doorknob throughout |
| “No extra people” | Only the current character appears in frame |
| “Follow storyboard cell 2” | State the starting posture, held object, and position shown in cell 2 |
| “This pass focuses on the text” | Preserve only the approved text surface, position, and content |
Finish with three checks: static information carried by the reference frame has been removed; every difference between start and end has an action that causes it; and the non-target facts most likely to drift have been written as preservation constraints. To derive start and end boundaries from a script, first read How to Write a Storyboard Script. For version-specific formatting for a particular video model, see Seedance Prompts.
FAQ
Should an image-to-video prompt repeat the character’s appearance?
No, when the reference image clearly establishes the face, hairstyle, clothing, and static composition. Keep local facts the motion depends on, such as which hand holds an object, the gaze target, the body’s position relative to the door or table, and motion-related occlusion.
The back of an object is not visible in the reference image. May the prompt invent it?
No. Include only back-side facts confirmed by the storyboard or asset information. Resolve unspecified color, text, or structure upstream in the asset or storyboard work. The prompt realizes established boundaries; it does not invent missing worldbuilding.
Why state the image-to-video endpoint separately?
The endpoint determines whether the shot completes its action and where the next shot can begin. Record body position, gaze, both hands, prop ownership, and the final visible composition. The next shot can then check continuity against the same fields.
Do it with the skill
In Claude Code, say: “Use /short-drama-video-prompts to turn the EP001 storyboard into shot-by-shot image-to-video prompts. Check each shot’s start state, action order, amplitude, camera behavior, and endpoint.” In Codex, write $short-drama-video-prompts. The skill reads existing storyboard and reference-frame declarations and produces text prompts plus end-state notes; it does not call a video-generation service.
Read the method: Reference subtraction, boundaries, and motion fields · Copyable body, text, and selective preservation · Contact order and action feasibility
About Drama SkillsThe short-drama-video-prompts skill on GitHub