Distinguish Timbre From Loudness and Delivery
Listen continuously across the join at normal speed, then to the source segments separately. Describe louder output, faster endings, emotional change, or an apparent different speaker. Emotion can vary within one voice.
If inconsistency is present in a segment, inspect synthesis first. If isolated segments sound similar but the export differs, inspect placement speed, processing, and consumed files. Without audio or listening, list unresolved checks rather than certifying a speaker change. Use the level guide for loudness.
Tie a Voice Name to Requests and Files
In the cited repository version, ordinary narration records the selected provider, model, voice, or reference configuration; segment caching checks text, settings, and WAV identity. The paths differ: MiMo uses its voice or an authorized reference, Fish a voice ID, and self-hosted IndexTTS an explicitly configured voice without local reference cloning.
An IndexTTS receipt records which voice was requested, not acoustic delivery of one person. A server may change the implementation behind an unchanged name without client detection. Invalidate affected tts_segments/*.cache.json records as documented to request again, preserving prior candidates and records rather than deleting all audio.
These are fixed-version workflow facts, not promises about every current provider feature or voice result.
Original Repair Plan: A Chosen Voice and One Old Segment
Original plan: three chosen Chinese narration segments occupy 0–4, 4–8, and 8–12 seconds. These are illustrative windows, not proven speech fits. Their texts are “她展开名单,” “她停在最后一行,” and “她把纸折了回去” (unfolds the list, stops at its last line, folds it back). The author chooses configured narrator A; no audio or listening exists yet.
| Segment | Version to check |
|---|---|
| 0–4 s | Current A request and candidate file. |
| 4–8 s | A manually imported old B candidate; renaming its record cannot make it A. |
| 8–12 s | Current A request and candidate file. |
Repair order:
- Preserve text, windows, and old B file; locate the middle segment’s provenance and request record.
- Synthesize that segment with the selected A configuration, not just a renamed file. If A’s server implementation changed, reassess regeneration for all affected segments.
- Actually listen through the first A, new middle, and final A segments, checking complete words and continuity. Reject unsuitable candidates; retain unresolved status without listening.
- Assemble the chosen actual files, verify that the export consumed them, and listen to the final join at normal speed.
Repository narration adoption must match explicitly supplied tts_meta, chosen text, and requested configuration. BOUND_TO_ADOPTION records consumption, not acoustic identity or naturalness. This paper plan has performed no synthesis, binding, or listening.
Do Not Substitute Level or Speed for Voice Repair
Protect exact text, audio roles, and windows. Level changes may smooth a join but cannot turn voice B into A; speed changes can create new delivery differences and are not a default timbre repair.
Use the window guide for overflow. With text preservation, resolve script, window, or explicit tempo budget instead of cutting final words. If multiple characters or narrators are intentional, confirm which differences belong.
Review Segments and Export With Specific Observations
| Check | Evidence |
|---|---|
| Requests and cache agree | Reduced configuration/version differences, not acoustic identity. |
| Candidate listening | Direct observations of these files, not every future generation. |
| Consumed-input record | Files and policy used in this export. |
| Normal-speed export listening | Actual continuity after placement and mixing. |
Record the segment, join, and audible change instead of “voice PASS.” Listening to only the middle cannot certify the whole video, and a new export needs its own check.
FAQ
Does the same voice name prove consistent timbre?
No. It is request configuration. File versions, server implementation, and actual sound still need checking; receipts and binding records do not replace listening.
Should I add effects if matching levels still sounds like a new speaker?
Return to source segments, requested voice, and consumed files. Avoid stacked processing hiding a wrong version, while retaining intentional delivery changes.
Do it with the skill
Use $video-voiceover to review voice consistency with problem segments, actual files, requested provider/voice, and cache records. Preserve text and windows; identify configuration or version differences and regeneration scope. Do not certify timbre without listening or substitute normalization for voice repair. Treat server implementation changes as a separate cache issue.
Read the method: Voice settings, segment cache, and requests · Named voices, cache limits, and receipt meaning · Narration inputs actually consumed