Describe What Sounds Mechanical
“Unnatural” is broad. Note an actual repeated emphasis, a pause after every subtitle fragment, a question delivered like an announcement, or a lively source that becomes rushed after placement. Different observations need different repairs.
Informational delivery may be steady; restrained narration need not sound surprised everywhere. Constant excitement or rising endings are not universal naturalness standards. With text alone, propose delivery and listening points rather than diagnosing or proving acoustic improvement.
Keep a Continuous Spoken Thought
The source editing method assigns narration a job—context, causal link, evidence-based interpretation, or transition—before writing a comprehensible thought. Subtitles may use several reading cues without turning the spoken passage into separately synthesized fragments.
“She arrives. She lifts the paper. She looks. She stops” may redundantly describe visible actions. If a connection is needed, relate the action to the change without inventing “she knows he is the killer” for excitement. Facts remain grounded in footage and selected context.
Retain useful quiet, dialogue, and reaction rather than filling every gap.
Check Whether This Path Consumes the Control
In the cited fixed version, the MiMo path places optional emotion in a natural-language user instruction and exact spoken words separately in assistant content. The emotion is not spoken text or voice verification.
That version’s Fish call does not pass this emotion field, so its presence in JSON does not establish use. Self-hosted IndexTTS sends voice and text only; nonempty emotion/style or nondefault rate/pitch fail explicitly. Remove unsupported settings or select another implementation rather than pretending they are understood.
These are adapter behaviors, not all provider products or a guarantee of instructed performance.
Original Plan: Keep the Words, Try Restrained Delivery
A paper plan uses two pictures, unfolding and folding back. The intervening look at the last line is an established story fact connected by narration in this version. Exact chosen text: “她把名单展开,看到最后一行,又折了回去。” There is no evidence she identified a killer or established authenticity. The provisional six-second window needs an actual fit check.
The chosen overall narration is restrained explanation, with an existing compatible style. Resolve a global style demanding exaggerated shouting before claiming a one-field change. Avoid subtitle-like chopped utterances. For the cited MiMo path, this is one candidate input without generated audio:
- start:
0 - end:
6 - narration: 她把名单展开,看到最后一行,又折了回去。
- emotion: 克制、疑虑;前半说明动作,末句收住,不夸张喊叫
The line means she unfolds the list, sees its last line, and folds it back. The delivery asks for restraint and doubt, explanatory first half and a held-back ending without shouting. That is a chosen narrator tone, not new character certainty; the actual example remains Chinese.
Comparison: retain text, voice, and window while changing this delivery instruction alone. Listen for emphasis, pauses, complete words, and the ending. Do not copy this input to a path rejecting emotion. Overflow needs a text-preserving decision about script, window, or explicit tempo budget rather than missing final words. A setting is not a listening result.
Check Delivery After Placement Too
Listen to the new source segment, then play the export at normal speed. Smooth source but rushed export calls for placement and tempo inspection; fragmented source calls for text and segmentation review. Check whether apparent improvement loses words, changes meaning, or covers essential original sound.
Record candidate, settings, text, window, and observations. Naturalness and impact remain listening judgments, not machine passes; a good instruction for one line is not every line’s template.
Use the pause guide for excessive waits without removing all breathing and quiet.
FAQ
Will emotion tags on every sentence sound more natural?
Not necessarily. Choose delivery for the passage’s role, check support, then listen. Steady explanation and restrained pauses may be appropriate.
Must voiceover use the same short splits as subtitles?
No. Subtitles serve reading while speech preserves thought. Multiple cues do not require separate synthesis for each.
Do it with the skill
Use $video-voiceover to review delivery with chosen text, window, segmentation, and actual observations. Preserve facts and exact lines; identify flow, supported controls, and downstream speed issues. Separate proposed text or window changes for author choice. Do not equate an emotion field with naturalness or send unsupported controls.
Read the method: Continuous spoken thought and audio roles · Delivery instructions, segments, and path differences · Explicit rejection of unsupported controls