Check whether the soundtrack still fits

Video Recap supports pairing independent picture with an adopted soundtrack. It fits a stable picture clock with a revised visual candidate while narration, source sound and music remain fixed.

Original example: a recap retains every shot duration while repairing the second shot’s picture. The editor wants the old mix. Another version removes three opening seconds and extends the ending; that changes timing, so old voice placement cannot simply be reused.

Record changed and retained elements, then check additions, deletions, speed and movement. Even equal total duration requires content review when speaking action or event order changes.

Select the two actual inputs

InputRecord
New pictureExample selection: Independently rendered picture.mp4 Check: Actual frame rate, start, duration and content
Retained soundExample selection: Selected complete AAC track from the old master Check: Which stream and which voice, source sound and music it contains
OutputExample selection: Paired candidate in a new directory Check: Old inputs and adopted version remain available

The cited version counts audio streams as a:N. Multiple tracks require explicit selection, not assumptions from filenames. The new picture’s audio and the donor’s video are ignored.

Do not regenerate TTS or add BGM when retaining a complete soundtrack; that would process it again.

Know what pairing preserves and cannot repair

The method is scripts/pair_media.py under video-assemble, driven by an explicit two-input plan. The cited implementation accepts suitable MP4-family H.264/HEVC zero-origin constant-rate picture and contiguous AAC audio, copying compressed streams instead of remixing.

It does not infer offset, speed changes, trimming, padding, normalization or shortest-stream truncation. Prepare suitable upstream inputs when format or timing fails; a failure is not successful preservation.

Endpoint tolerance checks container compatibility, not alignment between a spoken line and mouth movement. Stream and decode checks still require normal-speed viewing and complete listening.

Input-check-only planning records PLANNED and creates no paired video. Actual pairing that passes checks records PAIR_RENDERED. A FAILED record is not a successful candidate; early input errors also remain errors. Do not reuse a planned directory as the render directory; choose a new output directory and retain adopted files.

Complete pairing plan: choose the picture and adopted audio stream

Keep the example where only shot two’s picture changes and the complete clock stays fixed. This is a paper execution candidate: the author chooses /project/picture-v2.mp4 and the second audio stream, audio ordinal 1, from /project/adopted-final.mp4. Run only after checking the actual files, selected stream and internal event correspondence. These example files were not produced for this tutorial.

Save this complete plan as pair.json:

{"artifact": "media_pair", "schema_version": 1, "picture": {"path": "/project/picture-v2.mp4"}, "audio": {"path": "/project/adopted-final.mp4", "selected_stream": 1}}

Replace the paths with real files. Stream 1 is this example’s choice, not a universal selection. Verify that it contains the complete adopted narration, source sound and music. selected_stream counts audio streams only, not an overall index including video.

From the installed video-assemble skill directory, check the plan, then render into a different nonexistent directory:

python3 scripts/pair_media.py pair.json --output-dir pair-plan-check --plan-only

python3 scripts/pair_media.py pair.json --output-dir paired-candidate

Both directory names must be new. The first command creates no video; the second copies the picture and chosen audio streams. Renaming a planned directory does not make it rendered. There is no shift or tail-trimming field; unsuitable clocks require upstream input repair.

The output paired.mp4 has audio ordinal 0. The donor’s picture and the new picture’s own sound are ignored. Bind later subtitles to this new file and a:0, even though the donor used a:1.

Review the new picture in shot two, its internal speech/action correspondence and neighboring handoffs, then listen to the whole mix. Matching packets and clocks can still accompany mismatched content; see post-edit audio-video sync diagnosis. This example provides a complete plan and sequence, not executed commands or verified media.

Recheck subtitles after pairing

The paired output has one audio stream, a:0. Bind subtitles to that new container and output timeline rather than reusing the picture-only file or the donor’s old stream identity.

Check that picture uses the new version, sound comes completely from the selected retained track and subtitles match the current sound and placement. If subtitles also change, identify the lines instead of rerunning the recap and casually replacing its wording.

Preservation is an explicit input choice. Machine records identify consumed files and streams; listening and viewing establish editorial fit.

FAQ

Is equal duration enough for replacement?

No. Speech, action or reveals may have moved. Check actual starts, frame clocks and content correspondence; equal duration does not replace passage review.

Do it with the skill

Ask video-recap to change only specified picture, check timing, then use video-assemble to pair it with the selected complete soundtrack and rebind subtitles afterward.

Read the method: Pairing picture with retained audio · Audio processing and preservation modes · Explicit two-input pairing and result states

About Video Recap SkillsThe video-assemble skill on GitHub

All guides