Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
19 changes: 19 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,6 +4,25 @@ All notable changes to capcut-cli are documented here. The format follows [Keep

## [Unreleased]

## [0.29.0] — 2026-10-09

### Added

- `caption --words <file.json|->` builds captions from word timings produced by an external aligner, so Whisper does not have to be installed. The format is detected automatically: Whisper / whisper.cpp `segments[].words[]`, WhisperX `word_segments[]`, or a plain `{word|text|char, start, end}` array with times in seconds, `start_ms`/`end_ms` or `start_time`/`end_time` (per-character forced-aligner output). The result reports `words_format`, `words_skipped` and `source_words`. Entries without timing are skipped; entries that end before they start or go back in time are refused with `refused [words-invalid]` naming the entry. Grouping, CJK joining without spaces, `--script`, `--karaoke` and `--word-reveal` behave exactly as with Whisper words. `--words` cannot be combined with `--audio`, `--from-segment`, `--audio-stream`, `--ffmpeg-cmd` or `--whisper-*`.
- `batch <project> --plan <plan.json> < ops.jsonl` validates the operations under dry-run without touching the draft and writes a reviewable plan: `project`, `draft_sha256`, the normalized `operations`, `operations_sha256` (over key-sorted JSON) and a per-operation `preview`. `batch <project> --apply-plan <plan.json>` reads no stdin and applies the plan through the normal transactional batch. It refuses with `refused [plan-draft-changed]` when the draft changed since the plan was written, `refused [plan-tampered]` when the operations no longer match their hash, and `refused [plan-project-mismatch]` when the plan names another project. Without `--continue-on-error` a failing operation writes no plan.
- `render` (including `--dry-run`) reports a `fidelity` census of what the proxy leaves out compared with the app: `faithful`, `expected_duration_us`, the `checked` categories, and per dropped category a count, up to 20 segment ids and a hint (missing media, closed main-track gaps, uncomposited overlay tracks, unburned text, stickers, transitions, effects, filters, masks, keyframes, text/video animations, blend modes, chroma key, matting). `render --strict` refuses with `refused [render-unfaithful]` before anything is written.
- After a real render, `render` probes the written file with ffprobe (`--ffprobe-cmd` selects the binary) and reports `verification`: `{ verified, duration_us, expected_duration_us, drift_us, tolerance_us, within_tolerance }`, with a tolerance of one frame at the render fps. When ffprobe is missing or cannot read the file it reports `verified: false` with a reason. `render --verify` exits non-zero on drift beyond tolerance or an unverifiable file; the file stays on disk and the JSON result is still printed.
- `tts --lexicon <file.json>` applies pronunciation rules (`{"rules":[{"text","say","case_sensitive"?}]}` or a bare array) to the text handed to the TTS engine only: longest match first, at word boundaries for rule edges that are word characters (`SQL` never fires inside `MySQLite`; CJK rules match anywhere). Malformed lexicons refuse with `refused [lexicon-invalid]`, and equal-length rules that say different things at the same position refuse with `refused [lexicon-ambiguous]`, both before any engine runs. The result reports `lexicon: { rules, applied, spoken_text }` with offsets in UTF-16 code units into the trimmed text. Works with `--text` and `--text-file`.
- `lint` reports `segment-overlap` (error) when two segments on the same non-text track (video, audio, sticker, effect, filter …) overlap, naming the track, both segment ids and the overlap in µs. `lint --fix` repairs overlaps of at most one frame at the draft's fps (the independent-rounding overlap) by ending the earlier segment where the next begins, keeping its source duration in proportion to speed. Wider overlaps are reported, not fixed. Text tracks keep `caption-overlap`.
- `lint` reports `segment-offscreen` (warning) for video, photo, sticker and text segments that cannot be seen: the clip's box lies entirely outside the canvas (material size fitted into `canvas_config`, rotation allowed for), a scale axis is 0, or opacity is 0. Segments whose position, scale or alpha is keyframed are not judged on that property. No `--fix`.
- Every undo-history snapshot now records what made it: `.capcut-cli-history/<draft>.journal.jsonl` gets `{ index, time, command, argv, before_sha256, after_sha256 }` per write (library callers record `command: null`), trimmed together with the snapshots. `restore --list` shows `command`, `argv` and `time` per step; older snapshots without a journal line list with nulls. A torn or unreadable journal never breaks a write or a restore.
- Whisper-free caption routes are documented in `docs/quickstart.zh-CN.md` (不装 Whisper 的字幕路线) and `examples/short-video-narration.md`: the app's own auto captions followed by `restyle` / `export-srt`, `import-srt`, and `caption --words`.

### Changed

- `doctor`'s Whisper hint names `caption --words` and `import-srt` as routes that work without Whisper. The check's status is unchanged.
- `lint --frame-grid --fix` derives repaired source durations through the same helper as the new overlap repair: at speed 1 nothing changes; on clips at other speeds whose source duration matched target × speed (within 1%), the source duration now follows the snapped target duration instead of being left as it was.

## [0.28.0] — 2026-10-03

### Added
Expand Down
3 changes: 2 additions & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -105,9 +105,10 @@ The host reads a draft and passes its JSON as tool input. The component itself h

## Release notes

> **New in v0.29.0:** captions from any aligner's word timings without Whisper (`caption --words`); reviewed batch plans bound to the draft (`batch --plan` / `--apply-plan`); a render fidelity census and output-duration check (`render --strict` / `--verify`); TTS pronunciation lexicons (`tts --lexicon`); `segment-overlap` and `segment-offscreen` lint checks; and `restore --list` naming the command behind each step. Full details in the [changelog](./CHANGELOG.md).

> **New in v0.28.0:** opt-in active-timeline selection; compilation into empty app-created projects; compact command discovery and selection by name for agents; Windows command-path quoting fixed in Python client v0.1.3; and Python client CI on Linux, macOS, and Windows. Full details in the [changelog](./CHANGELOG.md).

> **New in v0.27.0:** content-safe media replacement; automatic import registration after replacement and relink; recursive, ambiguity-safe relinking; exact fractional-second compile boundaries; operation preflight and failed-build cleanup; payload-bound queue IDs; canonical project locks; bounded queue results; and local OTIO file/relative references. Full details in the [changelog](./CHANGELOG.md).

For an existing nested project, use `capcut diagnose <project> --active-timeline` to inspect the selected document before editing. The opt-in follows the pointer on unverified builds and refuses invalid or conflicting selected documents; it does not bypass write guards.

Expand Down
3 changes: 2 additions & 1 deletion README.zh-CN.md
Original file line number Diff line number Diff line change
Expand Up @@ -82,9 +82,10 @@ Claude Code 也可以把它作为插件加载:

## 发布说明

> **v0.29.0 新增:** 用任意对齐工具的逐字/逐词时间戳生成字幕,无需安装 Whisper(`caption --words`);与草稿哈希绑定、可先审后执行的批量计划(`batch --plan` / `--apply-plan`);渲染保真度清单与输出时长校验(`render --strict` / `--verify`);TTS 发音词典(`tts --lexicon`);`segment-overlap` 与 `segment-offscreen` 两项 lint 检查;`restore --list` 显示每一步由哪条命令产生。完整说明见[更新日志](./CHANGELOG.md)。

> **v0.28.0 新增:** 显式选择活动时间线、向应用创建的空项目编译,以及精简的命令发现索引与按命令名筛选;Python 客户端 v0.1.3 修复 Windows 命令路径的引号解析,并在 Linux、macOS、Windows 上运行客户端 CI。详见 [更新日志](./CHANGELOG.md)。

> **v0.27.0 新增:** 按文件内容安全替换媒体;替换和重链接后自动更新媒体导入登记;递归搜索并报告同名歧义;精确处理小数秒时间边界;编译前校验操作并清理失败输出;将队列 ID 绑定到任务参数;统一项目锁;限制队列输出;以及解析 OTIO 的本地文件 URL 和相对路径。完整说明见[更新日志](./CHANGELOG.md)。

## 使用 capcut-cli 构建

Expand Down
83 changes: 78 additions & 5 deletions docs/command-reference.json
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
{
"name": "capcut-cli",
"version": "0.28.0",
"version": "0.29.0",
"schema_version": 2,
"description": "Edit CapCut/JianYing draft_content.json directly. JSON in, JSON out.",
"global_flags": [
Expand Down Expand Up @@ -1405,7 +1405,7 @@
{
"name": "tts",
"summary": "Synthesize a voiceover from text via a local TTS command (--tts-cmd) and add it as an audio segment.",
"usage": "capcut tts <project> [start] [duration] (--text <string> | --text-file <path>) --tts-cmd <template> [options]",
"usage": "capcut tts <project> [start] [duration] (--text <string> | --text-file <path>) --tts-cmd <template> [--lexicon <file>] [options]",
"positionals": [
{
"name": "project",
Expand All @@ -1417,6 +1417,11 @@
"type": "string",
"required": true
},
{
"name": "file",
"type": "path",
"required": false
},
{
"name": "start",
"type": "time",
Expand Down Expand Up @@ -1456,6 +1461,15 @@
"required": false,
"description": "TTS command template, run without a shell: {out} (required) is replaced with the .wav path the tool must write, {text} with the text as one argument; without {text} the text is piped to stdin."
},
{
"name": "lexicon",
"flags": [
"--lexicon"
],
"type": "path",
"required": false,
"description": "Pronunciation lexicon JSON ({\"rules\": [{\"text\", \"say\", \"case_sensitive\"?}]} or a bare array) applied to the spoken text only: longest match first, at word boundaries (CJK rules match anywhere). Equal-length rules that say different things refuse [lexicon-ambiguous]; lexicon.applied offsets are UTF-16 code unit indices into the trimmed text."
},
{
"name": "volume",
"flags": [
Expand Down Expand Up @@ -3588,12 +3602,17 @@
{
"name": "batch",
"summary": "Run multiple edits from stdin (JSONL), one file write.",
"usage": "capcut batch <project> [--continue-on-error] < operations.jsonl",
"usage": "capcut batch <project> [--continue-on-error] [--plan <plan.json>] < operations.jsonl | capcut batch <project> --apply-plan <plan.json>",
"positionals": [
{
"name": "project",
"type": "path",
"required": true
},
{
"name": "plan.json",
"type": "string",
"required": false
}
],
"options": [
Expand All @@ -3606,6 +3625,24 @@
"required": false,
"description": "Commit only successful operations and exit 1 if any fail."
},
{
"name": "plan",
"flags": [
"--plan"
],
"type": "path",
"required": false,
"description": "Validate the stdin operations as a real run would and write a reviewable plan file (draft and operations sha256, per-operation preview) instead of changing the draft."
},
{
"name": "apply_plan",
"flags": [
"--apply-plan"
],
"type": "path",
"required": false,
"description": "Apply a plan written by --plan (no stdin). Refused if the draft changed since the plan or its operations were edited."
},
{
"name": "active_timeline",
"flags": [
Expand Down Expand Up @@ -4343,7 +4380,7 @@
{
"name": "caption",
"summary": "Transcribe audio via whisper into real caption-track segments.",
"usage": "capcut caption <project> (--audio <path> | --from-segment <id>) [options]",
"usage": "capcut caption <project> (--audio <path> | --from-segment <id> | --words <file.json|->) [options]",
"positionals": [
{
"name": "project",
Expand Down Expand Up @@ -4375,6 +4412,15 @@
"required": false,
"description": "Audio input."
},
{
"name": "words",
"flags": [
"--words"
],
"type": "path",
"required": false,
"description": "Word timings from an external aligner instead of running Whisper (path or - for stdin): Whisper/whisper.cpp segments[].words[], WhisperX word_segments[], or an array of {word|text|char, start, end} (seconds; start_ms/end_ms and start_time/end_time also read). The result reports words_format and words_skipped. Not combinable with --audio, --from-segment, --audio-stream or --whisper-*."
},
{
"name": "audio_stream",
"flags": [
Expand Down Expand Up @@ -5956,7 +6002,7 @@
],
"type": "boolean",
"required": false,
"description": "List snapshots."
"description": "List snapshots, newest first, with the command that made each write."
},
{
"name": "active_timeline",
Expand Down Expand Up @@ -6584,6 +6630,33 @@
"required": false,
"description": "Stream ffmpeg's progress to stderr instead of buffering it."
},
{
"name": "strict",
"flags": [
"--strict"
],
"type": "boolean",
"required": false,
"description": "Refuse (refused [render-unfaithful], nothing rendered) when the fidelity census finds anything the proxy drops: transitions, effects, filters, masks, keyframes, uncomposited tracks, stickers, text, animations, blend modes, chroma, matting, gaps or missing media."
},
{
"name": "verify",
"flags": [
"--verify"
],
"type": "boolean",
"required": false,
"description": "Probe the written file with ffprobe and exit non-zero when its duration drifts from the draft's by more than one frame, or when it cannot be probed. The file is kept."
},
{
"name": "ffprobe_cmd",
"flags": [
"--ffprobe-cmd"
],
"type": "path",
"required": false,
"description": "ffprobe binary for output verification."
},
{
"name": "active_timeline",
"flags": [
Expand Down
6 changes: 3 additions & 3 deletions docs/command-reference.md
Original file line number Diff line number Diff line change
Expand Up @@ -27,7 +27,7 @@
| `add-audio` | `capcut add-audio <project> <file-or-url> <start> [duration] [options]` | yes | Add a local or Wikimedia audio file on an audio track. |
| `add-video` | `capcut add-video <project> <file-or-url> <start> [duration] [options]` | yes | Add a local or Wikimedia video/image on a video track. |
| `add-text` | `capcut add-text <project> <start> <duration> <text> [options]` | yes | Add a text segment with font/color/position options. |
| `tts` | `capcut tts <project> [start] [duration] (--text <string> \| --text-file <path>) --tts-cmd <template> [options]` | yes | Synthesize a voiceover from text via a local TTS command (--tts-cmd) and add it as an audio segment. |
| `tts` | `capcut tts <project> [start] [duration] (--text <string> \| --text-file <path>) --tts-cmd <template> [--lexicon <file>] [options]` | yes | Synthesize a voiceover from text via a local TTS command (--tts-cmd) and add it as an audio segment. |
| `crop` | `capcut crop <project> <segment-id> [--ratio <r> \| --rect <x,y,w,h> \| --reset]` | yes | Read or set a video/photo segment's source-material crop (--ratio preset, --rect x,y,w,h, or --reset). |
| `cut` | `capcut cut <project> <start> <end> --out <path>` | yes | Extract a time range into a new standalone draft. |
| `duplicate` | `capcut duplicate <project> <segment-id> [--track <track-name>] [--new-track]` | yes | Duplicate a segment at its same timeline position onto a track above the source. |
Expand All @@ -51,11 +51,11 @@
| `apply-template` | `capcut apply-template <project> <template> <start> <duration> [text] [options]` | yes | Stamp a template into a project with new timing/text. |
| `make-preset` | `capcut make-preset <project> <text-segment-id> --out <preset.json>` | no | Extract a text segment's styling as a reusable preset JSON (apply via --preset). |
| `templates` | `capcut templates <project>` | no | List bundled reusable templates. |
| `batch` | `capcut batch <project> [--continue-on-error] < operations.jsonl` | yes | Run multiple edits from stdin (JSONL), one file write. |
| `batch` | `capcut batch <project> [--continue-on-error] [--plan <plan.json>] < operations.jsonl \| capcut batch <project> --apply-plan <plan.json>` | yes | Run multiple edits from stdin (JSONL), one file write. |
| `import-srt` | `capcut import-srt <project> <srt-or-> [options]` | yes | Import an SRT file/stdin as one text segment per cue. |
| `import-ass` | `capcut import-ass <project> <ass-or-> [options]` | yes | Import an ASS/SSA subtitle file as text segments, keeping inline overrides as per-range styles. |
| `text-ranges` | `capcut text-ranges <project> <id> --styles <json-or-@file>` | yes | Apply byte-accurate multi-style ranges to a text segment. |
| `caption` | `capcut caption <project> (--audio <path> \| --from-segment <id>) [options]` | yes | Transcribe audio via whisper into real caption-track segments. |
| `caption` | `capcut caption <project> (--audio <path> \| --from-segment <id> \| --words <file.json\|->) [options]` | yes | Transcribe audio via whisper into real caption-track segments. |
| `translate` | `capcut translate <project> --to <language> --out <path> [options]` | yes | Clone a draft into another language via the Anthropic API. |
| `migrate` | `capcut migrate <project> (--from <version> --to <version> \| --like <project> \| --from-store)` | yes | Apply known schema migrations across version boundaries. |
| `add-sfx` | `capcut add-sfx <project> <slug> <start> <duration> [options]` | yes | Add a sound effect on a dedicated track. |
Expand Down
Loading
Loading