Video workflow · 视频工作流
Video Recap Skills
Clip any video into a narrated Chinese recap — scene detection, script, voice-over, cut, subtitles — with a CapCut draft export.
把任意视频剪成中文解说视频——场景检测、写稿、配音、剪辑、字幕——并导出剪映草稿。
505 GitHub stars · 6 skills · MIT license · Format: Video workflow
Source on GitHub
English README
npx skills add zenstory-ai/video-recap-skills -y -g
What it is
Video Recap Skills is six independent skills plus an orchestrator for Claude Code that turn a source video into a narrated recap. It understands the footage (scene detection, ASR, a vision-language model), writes a director/edit/narration script, generates the voice-over, cuts and assembles with subtitles, and can export a CapCut (剪映) draft for hand-finishing. It runs on ffmpeg and one Xiaomi MiMo API key — no GPU, no model download.
它是什么
Video Recap Skills 是 6 个独立 skill 加一个编排器(Claude Code),把源视频做成解说视频:理解画面(场景检测、ASR、视觉语言模型)、写导演/剪辑/解说稿、生成配音、剪辑与字幕合成,并可导出剪映草稿手工精修。依赖 ffmpeg 和一个小米 MiMo API key——不需要 GPU,也不用下载模型。
Who it is for · 适合谁
- Recap and commentary channel creators
- Editors who finish in CapCut
- Teams turning long footage into short narrated cuts
- 影视解说 / 说剧类创作者
- 在剪映里精修的剪辑师
- 把长素材做成短解说的团队
How it works · 流程
- Understand Scene detection + ASR + VLM; optional background research so the model knows who is who.
- Script Director, editor and narrator passes write the recap script against the timeline.
- Voice Text-to-speech narration.
- Cut & assemble Audio mix, subtitles, render; or 'cut first, narrate after' mode.
- Export CapCut draft for manual finishing.
- 理解 场景检测 + ASR + VLM;可选背景调研,让模型认得人物。
- 写稿 导演 / 剪辑 / 解说三道稿子对着时间线写。
- 配音 文本转语音。
- 剪辑合成 混音、字幕、渲染;或「先剪后配」模式。
- 导出 剪映草稿,方便手工精修。
What makes it different · 有什么不同
- Local pipeline: ffmpeg + one API key, no GPU.
- CapCut draft export instead of a locked final render.
- Worked example shipped in the repo.
- 本地流水线:ffmpeg + 一个 API key,无需 GPU。
- 导出剪映草稿而不是锁死的成片。
- 仓库附完整示例。
Terms it uses
解说视频 剪映草稿 先剪后配