Distinguish translated dubbing from commentary

The English-to-Chinese mode translates the original speaker’s content and places sentence-level speech back on the source timeline. It does not lower the English track and add a rewritten commentary. Summarizing or deleting repetition belongs to a different editorial task.

Check speaker count and whether important music or ambient sound must survive. The pinned version supports one speaker and replaces the whole track; it does not separate music. Material with important location sound needs another audio plan, not an assumption that changing speech preserves everything else.

Verify transcription first. A mistaken name, quantity, negation or operating order remains wrong even in fluent Chinese. Listen to doubtful words before translating; synthesis cannot repair an upstream misunderstanding.

Preserve meaning, repetition and source windows

This original demonstration is a planning example, not a generated dub.

Source windowRecord
0–3 secondsEnglish: Keep the lid open until the light goes out. Chinese translation: 灯熄灭前,别合上盖子。 Check: Preserve the condition and negation
3–5 secondsEnglish: Not yet. Chinese translation: 还不行。 Check: Retain the repeated warning
5–8 secondsEnglish: Now close it, then press the blue button. Chinese translation: 现在合盖,再按蓝色按钮。 Check: Preserve order and color

Reuse the original sentence windows without overlap. Natural Chinese can reorder phrasing, but must not reverse the instruction or announce the next action early.

Record original text, approved translation, start, end, duration, doubts and listening result. Mark untested entries as pending, not synchronized.

Preserve meaning before judging spoken duration

The version uses roughly five Chinese characters per second as a drafting reference. Numbers, English names, pauses and voice affect actual duration, so it is not a universal ceiling. Remove translation-induced redundancy without dropping source information or merging sentences.

A verbose phrase meaning “during the period in which the indicator remains illuminated, do not perform closure” can become “while the light is on, do not close it.” That simplifies phrasing; removing the warning changes content.

The implementation anchors each line at its start and computes room using the earlier of that line’s end and the next start, with a 0.4-second minimum. It does not wait only for overlap with the next line. A WAV within room and tolerance retains its pace; an overlong one is compressed locally.

The cited version caps compression at 2×. If that remains insufficient, it trims the compressed audio to room and fades the tail. A fitted duration can therefore lose final words. Export success or matching duration is not complete-sentence proof.

Paper example: a five-second raw WAV in two seconds of room needs 2.5×. The cap first produces roughly 2.5 seconds, then trimming reduces it to two; the lost half-second may contain necessary content. No audio was generated or listened to here; actual lost words require listening.

If the complete meaning still cannot fit, revise the specific translation or timing plan and compare raw with fitted audio. Ordinary narration’s approved-text policy and dub are separate paths; a protection request does not automatically make dub block every tail trim.

Review against source speech, actions and joins

Check each sentence for lost conditions, complete delivery and correspondence with the action. Then play neighboring sentences together for overlap, awkward gaps or abrupt delivery changes.

For faulty transcription, repair the source text and translation first. For one overlong sentence, locate it rather than accelerating the entire video. If background music disappears, revisit whole-track replacement instead of diagnosing only a volume setting.

Successful export does not prove faithful translation or natural delivery. Keep the verified sentence table and listening notes so revisions can stay local. This guide concerns faithful cross-language dubbing; fitting an existing Chinese script to a picture edit is a separate question.

Dub’s input gate blocks empty, overlapping, and out-of-range lines and records text/timing risks before synthesis. It proves neither generated pronunciation nor content after trimming. Listen to conditions, negatives, and the final syllable, then check the picture.

FAQ

Can this mode directly dub a multi-speaker interview?

The cited version does not support that. It handles one speaker, sentence-level synthesis and whole-track replacement. Multiple speaker identities, voices and retained location audio need a different workflow.

Do it with the skill

Use video-recap for Chinese dubbing of this single-speaker English video. Verify transcription, preserve sentence meanings and source windows, and flag duration conflicts. First establish the effect of whole-track replacement on music and ambience.

Read the method: Sentence-level English-to-Chinese dubbing and scope · Per-line fit, compression cap, and tail trimming · Ordinary narration protection and dub boundaries

About Video Recap SkillsThe video-recap skill on GitHub

All guides