Make the sound decision explicit

No voiceover does not mean no sound design. Dialogue, action sound, ambience, and silence can carry content. Decide what to keep before choosing assembly.

NeedActual route
Adjust source level or add selected musicsource-mix: creates no narration, but processes sound according to settings, including the configured final loudness or limiter stage.
Preserve an adopted encoded audio trackadopted-packet-copy: copies a supported AAC stream and verifies its packets and clock without mixing.
Pair new picture with a complete mix in another fileUse the separate pairing method; do not substitute the new picture’s source audio for that mix.

source-mix is not frozen audio. “No voiceover” does not mean “unchanged soundtrack.” Copying has narrower input support: inspect codec, selected stream, and picture/audio intervals rather than relying on a filename extension.

Give the operator an unambiguous request

This fillable request is an operation brief, not a runnable script:

Input and adopted version: ___.

Add no voiceover; do not use old TTS as new audio.

Choose one: process source audio / preserve an adopted complete track.

Selected audio: ___; listen to identify the dialogue track if there are several.

Level or music decision: ___; add no BGM without a selection.

Picture range and subtitle requirements: ___; export a new candidate without overwriting the adopted delivery.

Report actual processing, then listen to speech, edit joins, and the ending.

The default entry is narration. Explicitly choose --audio-mode source-mix or --audio-mode adopted-packet-copy. --audio-stream-index is a zero-based audio-stream ordinal. Copying rejects TTS, configured BGM, and unsupported codecs. Report incompatible inputs rather than silently mixing and calling it a successful copy.

Original example: keep the sound of a repair demonstration

Suppose a repair demonstration has the teacher explain a stitch, demonstrate without speaking, then let a learner try. Existing sound contains the explanation and operation sounds; no new narration is needed.

To adjust audio and add already selected quiet music, use source-mix while preserving the teacher’s words and learner’s response. Listen afterward for masked operation sounds. The wordless demonstration is not an empty space that must be filled with music.

If an existing AAC track is already mixed and adopted, and only picture packaging should change, check copying eligibility. If it passes, keep that track without extra music or another loudness pass. Packet verification proves the copy contract, but full listening still checks content and whether the intended stream was selected.

This is a decision example, not an executed or listened-to result. If picture duration changes while another version’s sound must remain, resolve the intervals first rather than trimming an audio tail to force equality.

Original-sound example: retain explanation, action sounds and reply

Keep the hand-repair demonstration and select one concrete adopted-AAC copying route. These ten seconds, times and lines are original for this example, following the earlier mix-versus-copy decision. No file or listening approval exists here.

TimePicture and sound
0–3 secondsTeacher indicates the cloth and fully says, “Watch this stitch first.”
3–6 secondsContinuous needle passage and thread tightening, with original cloth and thread sounds.
6–10 seconds“Now try once.” Student: “Start here?” Teacher: “Yes.” The student begins.

Assume the ten-second picture is already edited and unchanged, with audio ordinal 0 holding the complete, previously heard and adopted AAC of these lines and action sounds. Only chosen visual packaging changes now, with no music or narration. Verify actual files, track and intervals against copying conditions; do not prefill the assumption as a pass.

Choose --audio-mode adopted-packet-copy and --audio-stream-index 0 explicitly, omit TTS and BGM configuration, and return a separate candidate. Copying performs no remixing, ducking, normalization or trimming. Packaging can re-encode picture, while copied audio still needs actual packet and timing verification. Existing captions and new visual text follow declared decisions instead of being removed automatically with narration absent.

Hear continuously from “Watch” through “Yes” and the subsequent action. Check the teacher’s ending near six seconds and an unobscured student reply. The speech-free 3–6 interval contains demonstration sound, without requiring music. Matching packets proves copying, not selection of the intended track or human content review.

If two seconds of waiting should later go, declare picture and sound changes and create another complete audio plan. Cutting this candidate to eight seconds cannot preserve the entire ten-second adopted track unchanged. The student starts trying; the excerpt does not prove mastery. Deliver according to its visible stopping point.

Check processing records and the actual soundtrack

Inspect assembly_qc.json for the actual mode, stream, and processing, then listen at normal speed to the opening, dialogue, joins, and ending. A video without new narration can still select the wrong track, retain old voiceover, cut a sentence, or process an adopted mix again.

For a JianYing draft, check its separate boundary: non-narration export currently supports selected audio stream 0 only. Other selections fail rather than silently reverting to 0. State subtitle and packaging requirements separately; absence of new TTS does not remove every old subtitle.

Source subtitles, narration subtitles, and new titles are different layers. Implement which to retain or remove according to this request. File or packet checks are technical evidence, not a listening or content-acceptance record.

FAQ

Does retaining source audio mean a lossless export?

Not necessarily. source-mix processes and encodes audio. Supported adopted-packet-copy preserves audio packets, while picture may still be re-encoded. Specify which stream preserves what and verify the result; do not call the entire video lossless.

Do it with the skill

Use video-assemble without narration. First determine whether the soundtrack is adopted, then choose source-mix or adopted-packet-copy. State the selected stream, music/level decisions, and input limits. Create no empty TTS, add no default BGM, export a new candidate, and identify listening still needed.

Read the method: Non-narration modes and copying boundaries · Inputs, default mode, and assembly entry · Pairing picture with an adopted track

About Video Recap SkillsThe video-assemble skill on GitHub

All guides