Find the first version that goes out of sync

Choose a moment where picture and sound really belong to the same event: a nearby door closing with its closing sound, or a clear utterance by the visible speaker. Added music accents, sound arriving from a distance and deliberate offscreen sound leads are not automatically coincidence anchors. Establish the edit’s intended relationship first.

Compare the source, the sound belonging to the locked edit, the mix candidate and the actual export. Record the file version, event, picture time, sound time and source location. Leave unobserved values unknown. If preview is suspect, compare the same export in another player before changing a file that may already be correct.

If the source is wrong too, address its footage or recording relationship first. If only the edit fails, inspect selected ranges and sound placement. For correct picture and sound with mistimed text, use subtitle synchronization checks. For incorrect generated mouth performance, use AI-video lip-sync diagnosis.

This is diagnosis of existing material. Rewriting the narration or recording every line again need not be the first step.

Compare several positions before choosing a repair

PatternCheck and repair
Several matching events in one passage are late by a similar amountCheck late placement or old timings. Change that passage’s position only when the evidence agrees and the complete sound remains available. One action is insufficient.
The start is close, but the gap growsCheck mixed versions, retiming of only one layer, and changed import or render clocks. Locate the difference before applying a global shift.
Sync changes suddenly after one recutInspect source boundaries and output positions around that cut, including an old bed still being used. Leave later correct passages in place.
One sound layer fits and another does notCheck dialogue, narration, effects and music independently; one late layer does not prove the whole mix is late.

These are editorial diagnoses, not automatic detection results. A door sound can be obscured or belong elsewhere; a waveform peak alone does not prove a shared source event. Check the complete line, the next action and another event after the cut until the repair scope has support.

Complete example: new picture In point, old source-audio In point

This fictional woodworking video is a paper example, not a media test. Its sources are chosen as zero-origin 25fps CFR footage meeting the source tool’s requirements; verify actual media in a real project. A door closes with its sound at source time 4.0 seconds. From 5.0 to 6.2 seconds, the speaker completes “Leave the key on the table.”

The finished picture lasts ten seconds: source time 1–7 maps to output 0–6, followed by another selected four-second shot. The failing candidate mistakenly uses source audio 0–6 at output 0–6. The second passage’s sound, continuous separate music and adopted narration placements are correct.

ItemExample record and decision
Correct first pictureSource 1–7 maps to output 0–6; the door closes at output 3.0, and the line belongs at 4.0–5.2.
Candidate first audioSource 0–6 was used; the door sound falls at output 4.0 and speech begins at 5.0. The source line’s final 0.2 seconds is also missing.
RepairRecover the complete source audio 1–7 and place it at output 0–6. This repairs event position and the lost ending. Moving the already shortened candidate alone is insufficient.
Adopted later materialKeep the second passage at output 6, with its correct sound, the continuous music and complete adopted narration. Do not shift these one second too.
Actual review still neededDoor correspondence, complete speech ending, the six-second cut, the second passage and the whole render; paper arithmetic is not listening approval.

At 25fps, the correct half-open source frame range is 25–175. Source frame 100, the door event, maps to output frame 75, or 3.0 seconds. On a 48kHz output clock, that is sample 144000. These values explain this example. Derive positions from the actual map for different rates or clocks.

Video Recap’s source/score producer prepares sound from explicitly selected source-frame ranges and output placements. Here only the first source range changes, but the complete bed plan must still contain the second passage and every explicit silent interval. Retain the chosen continuous music stem when it remains correct. Rebuild the affected bed and new mix-placement record, then use the existing full-sound assembly path for a new candidate. Unchanged text, WAVs and voice selection need no new TTS merely to move sound or repair a source range.

The bad candidate also lasts ten seconds, yet its first passage is mistimed and missing a word ending. After restoring the right source, review the door, whole sentence, cut and second passage at normal speed, then play the complete ten-second file. This review plan has not been executed.

Choose the path: preserving audio does not automatically synchronize it

PathCapability and boundary
Pair picture with adopted mixpair_media.py copies selected compressed streams and checks endpoints and packet clocks. It does not shift, retime, trim or perceptually align them. Internal events must still correspond.
Rebuild a source/score bedsource_score.py consumes explicit source-frame ranges and output placements, preparing source sound and continuous music. It does not infer dialogue, automatically repair drift or publish a final video.
Assemble adopted narration and bedThe full-sound path consumes chosen placements and gains while retaining complete sound. New output_start_sample values govern narration placement on this path; old generation windows are not automatically the new edit’s positions.

The current source-bed implementation checks CFR and actual packet clocks. It does not accept VFR, retiming or invented seconds-offset fields as supported operations. Prepare suitable inputs upstream or make an explicit editor revision, then check real correspondence rather than invent an unsupported option.

Adobe’s synchronization guide offers points such as In, Out, common timecode and markers. Identify which point represents the same event, then review the passage; a correct start does not prove all later events match.

YouTube upload troubleshooting separately addresses out-of-sync audio and video and advises checking track durations. In a local edit, also check individual events: equal duration does not establish content correspondence.

Review the new export

Play the actual new export at clear events before, within and after the change, then replay the passage and whole piece. Listen for complete line endings, repeated or missing speech at cuts and the adopted music continuity. Check that subtitles bind to the current sound version.

File and stream validation prove their tested mechanical relationships. Packet copy identifies retained audio; a bed receipt identifies source ranges; placement records identify positions. Perceived correspondence still requires playback of this render. Keep unreviewed positions pending.

Retain intentional sound leads, trails and dialogue crossing a picture cut. Synchronization does not require every sound to end neatly at a picture edit.

FAQ

Can I fix sync by moving all audio earlier?

Only when checked events support a consistent offset, complete sound is available and the move preserves other correct passages. Drift, wrong source ranges and lost endings need different repairs. This example recovers the source to repair position and completeness together.

Do aligned subtitles prove audio-video sync?

No. Captions can follow an incorrectly placed sound perfectly. Check caption-to-speech and sound-to-action correspondence separately in the new render; moving text does not repair misplaced audio.

Do it with the skill

In a host with Video Recap installed, use $video-recap only to diagnose this edit’s sound-picture mismatch. Use the supplied source, picture, mix and export versions, actual source-to-output map and observed event records. Find the first failing stage and smallest supported repair scope; leave unknown times unknown. Keep selected text, complete WAVs, voice and unaffected sound. List source ranges and new placements to verify. Do not rerecord, render or claim listening approval in this request.

Read the method: Source frame ranges and output sound beds · Complete adopted sound and explicit placements · Adopted sound after a recut · Packet copy and perceptual synchronization

About Video Recap SkillsThe video-recap skill on GitHub

All guides