First decide whether this is a join transient or another defect
Turn “there is noise” into a repeatable record: output file, exact time, defect span, whether the prior word is complete, whether the next word starts intact, and whether the same defect exists in source. Headphones can help locate it, but replay the complete sentence under normal listening conditions.
| Audible shape | Check first | Do not begin with |
|---|---|---|
| One very short click exactly at the join; both sources are otherwise clean | Non-continuous sources, abrupt edges and whether edge treatment reached the output | Re-recording everything or denoising the film |
| Final consonant, breath or release is missing | Source out-point and sentence/pause evidence | A fade that merely hides the missing word edge |
| Noise, reverberation or HVAC character changes after the cut | Location, source and ambience continuity | Calling a sustained change one click |
| A whole line crackles or clips | Recording, gain, mix and export | Fading only its beginning and end |
| Audio is clean but misses the visible action | Source intervals and placement | Treating synchronization as an edge transient |
A click already present in source is not repaired merely by moving one assembly boundary. If only one player preview clicks, replay the same export elsewhere before changing otherwise correct media.
Lock complete speech and true continuity before treating edges
The pinned video-cut workflow first avoids nearby source shot changes, then snaps boundaries toward reliable sentence ends or natural pauses. If ASR evidence still places an in/out point inside speech, it blocks. Complete sound has priority over cutting half a line for a cleaner picture boundary.
Separate two adjoining cases:
- A true continuous range from one source: The first clip ends exactly where the next begins, with no source time skipped. The implementation omits default edge fades here to avoid manufacturing a level dip. If this join still clicks, inspect source, encoding and the actual concat instead of assuming a longer fade is the only answer.
- Non-continuous clips or different sources: Abrupt waveform edges may meet. A short fade is a candidate, but it must preserve initial/final phonemes and meaningful silence and must be auditioned in output.
If “blue folder” is cut before the final consonant releases, extend the source through complete speech and the needed pause first. A smooth edge with an incomplete word is not a successful repair.
30 ms, 5 ms and 50 ms belong to specific implementations
These numbers describe source behavior; they are not one formula for every timeline.
| Pinned path | Implemented behavior | What it does not prove |
|---|---|---|
video-cut CLIP_JOIN_AUDIO_FADE_MS | Defaults to 30 ms; input must be finite and non-negative; each edge is capped at half the segment duration | That 30 ms is always inaudible or fixes every click |
| video-cut continuous-source join | A detected lossless continuation receives a 0 ms default fade on that edge | That all clips from one file are continuous; source intervals decide continuity |
| video-assemble narration edge | Fade length is constrained by measurable edge silence; without measurable silence it retains only a 5 ms anti-click ramp | Permission to cut off a word to fit; the implementation blocks when no complete safe window exists |
| short-drama-edit placed sound | Placed defaults to a 0.05-second tail fade, capped at half the sound, to avoid a click from a trimmed tail | A source-audio crossfade, execution of the prose sound note, or repair of distortion embedded in source |
Start with the current path’s actual default candidate and decide from what you hear. If changing an environment variable or external editor setting, record old and new values, affected joins and the new output rather than preserving only a UI screenshot.
Complete original diagnosis: two clips around a folder handoff
This is a paper exercise designed by the author. There are no real files, waveform measurements, renders or listening results. Two fictional community-centre sources surround a counter handoff. In source A, clerk Xu Zhou says, “The receipt is in the blue folder,” then slides it across the counter. In source B, volunteer Su Ning catches it, answers, “I have it,” and puts it in a transfer box. The story must preserve instruction, receipt and acknowledgement; “I have it” does not mean final delivery.
Suppose the first plan is:
| Clip | Designed source range and job |
|---|---|
| A | Source A 4.20–6.80 s: complete instruction; folder reaches the counter midpoint; Xu releases it |
| B | Source B 12.40–14.10 s: Su catches it, completes “I have it,” and puts it in the box |
| Output join | B follows A; imagine one click at output 2.60 s. Nothing was generated, so this example cannot say the click was detected or fixed |
Build three candidate diagnoses before choosing a parameter:
- If fictional source A itself clicks near 6.80 s, the defect belongs to the source tail or selected range. Compare an earlier/later legal out-point that still preserves release, or treat source audio. Do not claim a concat fade removed an embedded defect.
- If each source is clean and both speech/action units are complete, while the anomaly exists only at the non-continuous output join, preserve their jobs and ranges and render a candidate with this video-cut version’s default edge treatment. Listen at normal speed from “blue folder” through the catch, then inspect the join. 30 ms is an implementation fact; the result remains untested.
- If B carries a hollower counter ambience throughout, the defect is a sound-field change. Route to ambience continuity: record space, noise and reverberation and compare a proper bed or external cross-treatment. Lengthening the edge fade from 30 ms does not establish a unified sound field.
Complete-speech branch: Suppose an actual review later finds that A’s 6.80 s out-point clips the final release in “folder,” while 6.96 s contains the full release and a short pause. Extend candidate A to the real 6.96 s first, then check that the release and B’s catch do not duplicate the handoff. If action ranges conflict, choose coverage that preserves both speech and the one handoff rather than fading away the phoneme. Here 6.96 s remains an invented teaching value, not a number for other videos.
Only a new output can support the final verdict: complete words, one handoff, no audible click, appropriate ambience and a bounded acknowledgement that Su received the folder. Mark anything not watched or heard as pending.
Copy this join-repair record
Change only the evidenced scope:
Current output: Path, version/hash and join time.
Previous source: File, source in/out and last complete phoneme/action.
Next source: File, source in/out and first complete phoneme/action.
Defect shape: One click / clipped word / sustained ambience change / passage-wide distortion / sync offset.
Source comparison: Whether each source contains the same anomaly by itself.
Continuity: Same source with exactly adjoining intervals, or unknown.
This change: Boundary moved, phoneme restored, pinned-tool default used, or explicit external treatment.
Keep unchanged: Picture jobs, full lines, captions, narration, score or other accepted tracks.
New-output review: Complete sentence at normal speed, both sides of the join, one action, ambience, captions and sync; mark unheard items pending.
For video-cut, retain actual source seconds and reasons in clip_plan.json; do not patch only the export and leave a stale plan. If narration was authored against the old output timeline, recheck it against the new edited_source.mp4 rather than assuming later positions remain correct.
Failure repairs: one fade does not solve every audio defect
| Failed approach | Risk | Smaller repair |
|---|---|---|
| Lengthen crossfades at every cut | A continuous-source join dips; consonants and short effects weaken | Establish continuity and defect location; treat only evidenced edges |
| Fade a missing word ending | The loss remains, merely quieter | Restore the complete source ending and needed pause, then choose an action join |
| Treat an ambience jump as one click | The room still changes throughout the next passage | Work on the ambience layer, documenting source, distance, reverberation and actual external treatment |
| Repair only the edges of a clipped passage | Distortion in the middle remains | Compare recording, mix and export and repair the first failing layer |
| Stop when a tool reports no issue | Successful processing is not an audition | Listen to the new join, full sentence and passage, preserving untested items |
| Trim a TTS word to fit | Meaning and intelligibility are damaged | Shorten text, move the window or block; the pinned narration path does not call an incomplete tail safe fitting |
FAQ
Does a longer fade always remove a click more safely?
No. A longer fade can weaken phonemes, short effects or a continuous join. First prove the defect is edge-only, speech is complete and the clips are not a true continuation, then audition the candidate.
Is video-cut’s 30 ms a universal recommendation?
No. It is the pinned source version’s default, further limited by clip duration and the continuous-source exemption. It is not a standard for every editor, source or sound.
Do it with the skill
Use $video-recap to inspect the marked audio join. Compare the actual output with both source segments, checking complete phonemes, continuous-source status and defect duration. Preserve source intervals and reasons in the clip plan, make the smallest repair candidate, and render a new output. Do not lengthen fades at every join; report only after listening to the new file.
Read the method: video-cut joins, speech boundaries and audio priority · 30 ms default and parameter validation · Continuous-join exemption and edge-fade implementation · Narration edge silence and 5 ms anti-click fallback · Short-drama effect-tail fade and mixing boundaries