Use a two-column check to bind now or hold
| Check | Binding decision |
|---|---|
| Has the voice been selected? | Auditioned, accepted, and registered: continue. Still marked pending or represented only by adjectives: hold. |
| Is the file real and readable? | Project-relative path opens: continue. Filename, prose description, or memory alone: hold. |
| Does the target accept reference audio? | Current target profile explicitly accepts it: continue. Capability absent or unknown: retain the voice record and omit the binding. |
| Who speaks in this shot? | Bind only the character who delivers the scripted line. Do not bind a silent listener or a visible nonspeaker. |
| What role does this reference serve? | This example uses timbre only. A separate accent, pace, or age reference receives its own reviewed binding. |
| What must be excluded? | Sample words, delivery, emotion, pace, pauses, recording space, background sound, and other speakers. |
| Where does the current line come from? | Copy it exactly from the current script. If several characters have spoken the same line elsewhere, name this shot’s speaker. |
This step follows voice selection. A written description such as calm, young, or rough cannot serve as an audio reference, and an unselected candidate cannot be attached first and decided later. Finish auditioning, usage authorization, and project registration before bringing the selected real path into the video prompt.
Copy the path and number image and audio inputs separately
Keep the project’s native video-prompt field order. The literal 参考音频 line follows the literal 输入参考图 line. Image ordering starts at one, and audio ordering starts at one independently. The first audio remains 音频1 even after three images. Copy the registered project path exactly so that later shots cannot switch files silently.
| Native field | Value in this example |
|---|---|
输入参考图 | - 输入参考图:REF-XC(顺序:1)· 输入/角色/许澄-工作服.png《许澄工作服》;REF-LX(顺序:2)· 输入/角色/罗弦-工作服.png《罗弦工作服》;REF-ROOM(顺序:3)· 输入/地点/修表间.png《修表间》 |
参考音频 | - 参考音频:REF-VOICE-XC(顺序:1)· 输入/声音参考/许澄-音色-v3.wav《许澄已选音色》(用途:音色;角色:许澄;控制:音色质感、共鸣、年龄与体量印象;不得控制:台词、语气、情绪、语速、停顿、录音空间、背景声、其他说话人) |
| Current speaker | 许澄; he is the only person speaking in this shot. |
| Exact current line | 先把旧标签留在盒里,明早一起核对。 |
| Silent listener | 罗弦; she reacts after looking at the old label and receives no voice reference. |
The native 顺序:1 value identifies the first audio input; it does not continue after the three image slots. The audio line stays outside 输入参考图, and the storyboard receives no new audio binding. The surrounding English explains the record while the copy-ready schema remains in its required literal Chinese form.
Complete original example: keep the old label in the repair room
Tomorrow Morning Check is an original paper shot created for this article. No audio was uploaded, no video was generated, and no lip movement or resulting sound was reviewed. This worked example belongs to one Chinese-script project. Its registered character names are 许澄 and 罗弦; the surrounding English explanation glosses them as Xu Cheng and Luo Xian. Xu Cheng runs a watch-repair shop, and Luo Xian is an apprentice who has just begun handling jobs independently. Luo has repaired an unclaimed old watch and wants to discard the faded label inside its box so the delivery looks tidy. Xu knows the label could help an older visitor identify the object the next morning. He also sees that Luo expects blame for leaving clutter behind.
The teaching record treats 输入/声音参考/许澄-音色-v3.wav as a selected placeholder path. Its designed audition words are 雾港的钟声过了三遍, delivered brightly with audition-room reverberation. It supplies 许澄’s timbre only. Its words, bright delivery, emotion, pace, pauses, and room sound stay outside the current shot. The current script line is 先把旧标签留在盒里,明早一起核对。 罗弦 says nothing. She begins with one corner of the label between her fingers, then lays it flat in the box after hearing him.
For reader explanation only, the current line means “Leave the old label in the box. We’ll check it together tomorrow morning.” This English translation is not a second script line and does not belong in the native record or executable prompt.
| Shot element | Established content |
|---|---|
| Start | Watch-repair bench; old watch in an open wooden box; 罗弦 holds one corner of the faded label; 许澄 stands beside the bench looking toward it |
| Image references | Image one fixes 许澄’s identity and workwear; image two fixes 罗弦’s identity and workwear; image three fixes the bench, box, lamp, and tools on the back wall |
| Single action | After 许澄 completes the current line, 罗弦 lays the label flat in the box; she has no dialogue |
| Sound | 许澄 alone delivers 先把旧标签留在盒里,明早一起核对。; established quiet ticking continues underneath; no audition words, audition-room reverberation, or added voice |
| End | Label lies beside the old watch; 罗弦’s fingers leave it and she looks toward 许澄; 许澄 remains beside the bench |
| Voice binding | 音频1 supplies 许澄’s timbre only; 罗弦 is a silent listener and receives no binding |
The current target profile in this paper project has already established 音频1 as its token for the first reference-audio input. The following executable prompt keeps the registered Chinese names and exact Chinese script line. It makes no claim about an unnamed vendor format:
Keep the watch-repair room from 图片3, 许澄’s identity and workwear from 图片1, and 罗弦’s identity and workwear from 图片2. 许澄 stands beside the bench and looks toward the old label in 罗弦’s hand. 音频1 supplies only 许澄’s timbral texture, resonance, and age and body impression. With the calm, clear delivery established for this scene, 许澄 says in Mandarin Chinese, “先把旧标签留在盒里,明早一起核对。” Preserve that Chinese line exactly. 罗弦 remains silent. After the line, she lays the label flat in the wooden box, releases it, and looks toward 许澄. Quiet clock ticking continues through the shot. Exclude the audition words, audition delivery, audition emotion, audition pace, audition pauses, recording space, background sound, and every other speaker from the voice reference. End with the label beside the old watch and preserve both characters’ positions and every bench prop.
The binding keeps both motives legible. Luo wants a clean handoff, while Xu chooses to preserve old evidence for one more night. The result changes what they can check tomorrow and turns immediate disposal into a shared decision.
Let the current script control words and performance
The voice reference carries identity: timbral texture, resonance, and age or body impression. The current script and shot performance carry the event: exact words, speaker, delivery, emotion, emphasis, pace, pauses, volume, and present breathing. Separating these layers lets a character remain recognizable across scenes while changing expression during an argument, reassurance, concealment, or ordinary work.
Assign every audible feature in the sample to a clear owner. Selecting a timbre leaves sample words behind. A pleasing sample pace does not become the character’s permanent pace. When the recording contains room sound, an exclusion written in the prompt does not clean the audio file. Background noise, an extra speaker, or a usage-authorization problem returns to voice-asset work for cleaning, rerecording, or replacement, followed by an updated registered path.
When several people have spoken the same line elsewhere in the source scene, name the current speaker in the prompt body. Bind voices by actual speakers, not by the number of visible characters. Only Xu speaks in this example, so the shot has one audio input. Luo’s visible action carries her response while she remains silent.
What is complete after the binding record?
$short-drama-video-prompts can write and inspect the binding record for a shot: whether the path was copied, the registered character truly speaks, purpose and exclusions are complete, audio order is independent, and the current line appears once. It does not audition candidates, host audio, upload the file to a target service, start generation, or confirm that the resulting video retained the selected timbre.
Before submission, verify five items: the reference file remains readable; registered character and current speaker match; the target profile still accepts reference audio; the prompt contains only the current scripted line; and silent characters have no extra binding. After generation, listen to the real output for character recognizability, exact words, current performance direction, carried noise, and added voices, then inspect lip movement separately. This example contains only a paper binding record, so it reports no generation or listening result.
FAQ
Can I add a placeholder path before selecting the voice?
Keep the voice marked pending and do not invent a bindable path. Complete auditioning, authorization review, and project registration, then copy the real path into the shot where the character speaks.
Do both visible characters need voice bindings?
Bind only the person who actually speaks in the shot. A silent listener carries the response through expression, eyeline, or action and receives no audio binding.
Do it with the skill
Use $short-drama-video-prompts to bind the selected character timbre to this actual speaking shot. Copy the real registered voice path, add independently numbered reference audio after the image references, and limit its role to timbre while excluding sample words, delivery, emotion, pace, pauses, recording space, background sound, and other speakers. Preserve the current script exactly, bind only actual speakers, inspect the record, and stop before upload or generation.
Read the method: Character audio-reference binding in video prompts · Voice-reference media, scope, and registration
About Drama SkillsThe short-drama-video-prompts skill on GitHub