The Three H3 Fields: Timeline, Soundscape, and Score
These English field names belong to H3's structured prompt format; they are not a universal formula for every video model. Fill them as follows:
- **
integrated_multimodal_description**: Beginning with[Shot 1], describe framing and composition, the visible start state, observable action, speakers, verbatim dialogue, synchronized sound, and a checkable end state in playback order. Keep[Shot 1]even when the video contains one shot. - **
overall_soundscape**: In one continuous paragraph, summarize ambience, physical action sounds, and non-verbal human sounds across the video—rain, footsteps, a pot touching a counter, or breathing. Do not repeat dialogue here. - **
non_diegetic_music**: Describe only music that the audience hears and the characters do not, using instrumentation, tempo, rhythm, and changes in volume. A radio or live performance audible inside the scene remains on the main timeline. WriteN/Awhen no audience-only score is intended.
The sound-placement test is whether a cue must land on a particular action or vocal beat. Put time-critical diegetic sound beside that action on the main timeline. overall_soundscape summarizes ambience, physical sounds, and non-verbal human sounds across the clip; it may briefly recap an important cue, but it should not repeat dialogue.
Copyable skeleton:
integrated_multimodal_description: [Shot 1] [Style and framing]. [Visible start state]. [Action and visible result; add a timed in-scene sound cue beside its triggering action, if needed]. The [speaker description] (S1) says: <d>[Chinese] 逐字台词</d>. [Visible end state].
overall_soundscape: [Summarize clip-wide ambience, physical sounds, and non-verbal human sounds; do not repeat dialogue].
non_diegetic_music: N/A
If you want an audience-only score, replace the final line of that block with:
non_diegetic_music: [Audience-only score: instruments, tempo, rhythm, and dynamics]
This page follows the English field names and example style of H3's official structured guide. That describes the template; it does not claim English prompts work better or that every interface requires English. Give each speaker a stable ID such as (S1). Put the speaker's identity, voice, or action outside <d>; inside <d>, preserve the language tag and spoken words. Keep Chinese dialogue verbatim as <d>[Chinese] original line</d>. See the official base and keyframe prompt guide.
These are prompt structures and reading examples, not a video-task request JSON. Preserve field names, speaker IDs, and dialogue tags. The <d> markers are prompt text that still belongs inside the interface’s text input. The linked official guide at its fixed version was rechecked on September 30, 2026; this edit ran no video generation.
First Frames, First-and-Last Frames, and Full Reference Use Different Shapes
Choose the prompt shape from the inputs you will actually submit. Do not start from a template and decide whether to attach references afterward.
| Input | Record |
|---|---|
| No image, with text-to-video intentionally selected | Structure: Three fields; Writing focus: Establish the complete visible start, action path, and end in text |
| One first frame | Structure: First-frame alignment instruction + three fields; Writing focus: Continue forward from the pose, composition, and scene already present in the first frame |
| One first frame + one last frame | Structure: First/last-frame alignment instruction + three fields; Writing focus: Describe an observable continuous path between the two frames; arrive at the last frame without adding a new state afterward |
| One last frame only (if the interface supports it) | Structure: Last-frame alignment instruction + three fields; Writing focus: Begin from a plausible earlier state and converge on the supplied last frame |
| Ordinary character, location, or prop images; reference video or audio; or a first frame plus other reference assets | Structure: Six-section full-reference format; Writing focus: Define the labels and jobs of the actual assets before describing retention and the detailed timeline |
The official guide also defines last-frame-only L2VA: if your current interface accepts that input, add its last-frame alignment instruction and use the three fields to move from a plausible earlier state toward the supplied ending. Available modes differ by interface and API.
The full-reference sections are:
subject_definitions: ...
summary: ...
retention_analysis: ...
detailed_description: [Shot 1] ...
overall_soundscape: ...
non_diegetic_music: ...
The six sections form one complete prompt. They are not six separate generations, and detailed_description should not be submitted alone. Labels such as <Picture 1>, <Video 1>, and <Audio 1> must identify assets actually attached to this generation task; numbering within each asset type must match the submission order. Merely typing “use the character reference” or <Picture 1> cannot replace supplying the file. See the official full-reference prompt guide for the complete six-section rules. The UI or API you use may not expose every mode, so follow the inputs it actually accepts. Do not silently turn a missing-reference task into text-to-video.
The official video-generation API distinguishes media roles separately: a request with reference images, video, or audio cannot also contain first_frame / last_frame roles. For an opening image plus a character reference, supply them in reference mode and identify which image controls the opening composition in the prompt. An alignment sentence cannot replace an attachment. This table does not create a mode absent from your interface.
Original Example: A New Leaf in a Flower Shop After Rain
The following is a writing example, not a tested H3 result. It is a text-to-video, single-shot, locked-camera setup: after a rain shower, a neighborhood flower-shop worker notices a new leaf, moves its pot from a wet windowsill to a dry counter, and speaks one Chinese line.
integrated_multimodal_description: [Shot 1] Live-action, a locked medium shot inside a small neighborhood flower shop just after a rain shower. A shop worker in a plain green apron stands beside a rain-speckled front window. A small potted pothos rests on the wet windowsill in front of her. She notices one pale new leaf beneath the larger leaves and leans closer without touching it. She then cups the pot with both hands, lifts it straight up from the wet sill, carries it one step to the dry wooden counter on the right, and sets it down with a soft ceramic tap. Keeping both hands beside the pot, the worker with a clear, gentle voice (S1) looks at the new leaf and says: <d>[Chinese] 你也撑过这场雨了。</d> She ends facing the pot on the dry counter, her gaze still on the new leaf.
overall_soundscape: Rainwater drips at an uneven pace from the awning outside while distant tires pass over the wet street. As the pot is set down, its base makes one light tap on the wooden counter.
non_diegetic_music: N/A
The prompt carries one main change: the pot moves from the wet sill to the dry counter. Its start, action, and end are observable within one shot. The line appears once and is bound to the worker through (S1). Place the pot's light tap beside the setting-down action on the main timeline, then briefly summarize it in overall_soundscape; post-rain street ambience belongs only in that overall summary. Because no audience-only score is intended, the final field explicitly says N/A. If you convert this to first-frame or first-and-last-frame generation, attach the required image or images first, add the guide's alignment instruction, and make the body begin from the first frame or land precisely on the last frame.
Revise the Field That Contains the Problem
Inspect the text first, then validate it against an actual generation. These revisions do not guarantee that the model will follow every instruction.
| Problem | Revision |
|---|---|
Chinese dialogue has been translated or polished, or <d> contains delivery notes such as “softly” | Restore the original words and punctuation, keep [Chinese], and move delivery, action, and speaker identity outside <d> |
| A line exists, but the speaker is unclear | At the speaker's first vocal event, add an identifiable description and (S1); reuse that ID for later lines from the same speaker |
| A timed cue appears only in the overall soundscape, or ambience and dialogue are all pushed into the main timeline | Put time-critical in-scene sound beside its action in integrated_multimodal_description; let overall_soundscape summarize clip-wide ambience, physical sounds, and non-verbal human sounds without repeating dialogue; reserve non_diegetic_music for audience-only score |
| First and last frames are fixed, but the body keeps the character moving or puts the prop in a third position after the ending | Remove every action after the last-frame state; rewrite the middle as a continuous path from the visible first-frame state to the last-frame state |
The prompt names <Picture 2>, a character image, or reference audio that is absent from the task inputs | Attach the asset and verify the order within its asset type. If you no longer intend to use it, remove the label and choose the no-image, first-frame, first/last-frame, or full-reference structure again; text cannot pretend a file was supplied |
Revise only the field where you have located the problem. If the ending itself is ambiguous, define it in script-to-storyboard first. For general image-to-video questions involving hands, props, occlusion, and start/end differences, use image-to-video prompts. This page stays focused on H3's format, dialogue and sound layers, and the structure required by the actual inputs.
FAQ
Should a Hailuo H3 prompt be written in Chinese or English?
The copyable template on this page follows the English field names and example format in H3's official guide. Do not translate Chinese dialogue; preserve it verbatim inside <d>[Chinese] ...</d>. This describes the template, not a claim that English works better or that every interface requires it. If your interface accepts another language, follow its prompt requirements.
Which structure should I use for first-frame, last-frame, and reference inputs?
For one first frame, use its alignment instruction plus the three fields. For one last frame only—if your interface supports it—use the L2VA alignment instruction and describe a plausible path toward that ending. A first/last-frame pair also uses an alignment instruction plus the three fields. Ordinary reference images, video, or audio use the six-section full-reference format. Submit every asset and match labels to the order within each asset type.
Where do dialogue, ambience, and music go?
Put dialogue and any in-scene sound that must hit a specific action beat at that point in integrated_multimodal_description, or detailed_description for full reference. overall_soundscape summarizes clip-wide ambience, physical sounds, and non-verbal human sounds; it may briefly recap an important cue, but do not repeat dialogue. Reserve non_diegetic_music for music only the audience hears. Radio, live performance, or singing audible to the characters stays on the main timeline.
Do it with the skill
Prepare the approved single-shot storyboard, verbatim dialogue, and the first frame, last frame, or reference assets you will actually submit. Then use /short-drama-video-prompts (or $short-drama-video-prompts in Codex) and explicitly request the MiniMax H3 format. It selects the three-field or six-section full-reference shape from the real inputs and produces copyable prompt text; this step does not upload assets or run video generation.
Read the method: H3 format, input modes, and sound layers
About Drama SkillsThe short-drama-video-prompts skill on GitHub