Install
openclaw skills install @aiaaaa4/video-translate经确认把视频音频交给 OkFile/Fun-ASR,并用 Qwen 或当前 Agent 生成和全文审校双语字幕。
openclaw skills install @aiaaaa4/video-translate作者 / 工作流设计:AI落地第四声。本作者信息用于展示和来源识别,不添加额外授权限制。
这是一套面向本地录制视频的高质量字幕翻译工作流。OkFile + Fun-ASR 固定负责云端转写和词级时间戳;原语言字幕只能在当前 Fun-ASR 词跨度内校正识别内容,不能借用相邻段或替代边界,异常时整份参考源自动关闭。翻译前,当前 Agent 必须先通读完整源文,生成本视频专属的领域提示、术语、专名、歧义判断和翻译记忆。公开默认由 qwen-mt-plus 稳定初译;用户也可选择当前 Agent 编排模型直接翻译。随后 Agent 再次通读原文和译文,重新翻译歧义或错译内容并按语义重分段;确定性 QA 后再完成最终全文 QC,全部通过才导出双语 ASS/SRT。
快速开始:准备 OkFile API Key、阿里百炼 API Key、阿里工作空间 ID,并提供本地视频路径。用户直接提供音频时必须拒绝;组合工作流仍可在内部复用 video-download 写入 .work/input/ 的音频。AI 会从原文件名或媒体项目目录提取真实标题,去掉开头日期、结尾平台编码和扩展名,再按视频领域术语翻译为中文;最终 ASS/SRT 使用该中文净标题,不会使用“原版视频”等占位名。AI 仅在你明确确认视频路径、翻译模式、输出位置和外发处理同意后,才会读取本机 .env、上传处理音频并运行固定生产流程。
.work/input/ handoff files, the confirmed output directory, and the local .env required for the fixed providers.https://www.okfile.com and its temporary public URL is sent to Alibaba Fun-ASR. In qwen mode, subtitle text is also sent to Alibaba qwen-mt-plus; in Agent mode, transcript text is handled by the current Agent model service. No external processing starts without an explicit user confirmation..work/input/ handoff directory.首次使用时必须先运行固定问卷。工作流会本地准备音频、通过 OkFile 与 Fun-ASR 获取词级转写;当前 Agent 先通读完整源文并生成翻译上下文,再按用户选择由 qwen-mt-plus 或当前 Agent 完成初译,之后 Agent 对照原文完成全文重译审校、语义分段和最终 QC。
qwen-mt-plus,因为它稳定、性价比高,并支持术语、翻译记忆与缓存恢复;追求极致质量时可选择当前 Agent 编排模型直接翻译,但会消耗当前 Agent 的模型额度,耗时和质量取决于所选模型。若当前环境支持,推荐在 Codex 中使用 GPT-5.6。DASHSCOPE_API_KEY、ALIYUN_WORKSPACE_ID、OKFILE_TOKEN。环境缺失时,先说明用途并询问用户是否同意创建并打开本机 .env;只有得到明确同意后才执行 bash scripts/open_env_setup.sh --open。不得让用户在聊天中发送密钥。用户示例:
把 /Users/me/Desktop/lesson.mp4 翻译成中文字幕。保留原文,视频里的 PPT 文字也很重要。
Use this skill only for local recorded video. Reject user-selected audio files. Before every production run, read the full execution contract in full.
Keep this production stack fixed unless the user explicitly requests an engineering redesign and accepts revalidation:
.work/input/ directory or a same-basename audio-only download beside that video; otherwise extract compact audio locally with ffmpeg. Never accept a user-selected audio file as the workflow input..work/input/ contains one original-language SRT/VTT, require sufficient time overlap and lexical similarity, crop the reference to the current ASR word span, and use it only to correct SRC_DISPLAY and translation source text. Never copy an entire cross-boundary cue into a smaller ASR segment, replace SRC_RAW, or invent word timestamps. Disable an anomalous reference source and regenerate display text from Fun-ASR.domains, terms, and tm_list to qwen-mt-plus and bind its cache to the context hash. Agent mode: translate every hash-bound section using the same context and write validated receipts. In either mode, each ZH_i translates only its own SEG; never advance, delay, split, or distribute meaning across neighboring SEG blocks. Keep an incomplete source fragment equally incomplete until semantic review.ZH_i with its own source and at least the ±2 neighboring sources, and treat a better neighbor match or consecutive offset pattern as a blocker. Correct mistranslations and re-segment only in semantic review. If all reviewed content is unchanged, require an explicit no-change confirmation rather than trusting a bare passed receipt.SRC_RAW, run deterministic QA, then require final whole-document QC with the same cross-segment alignment check and fixed spot checks.--localized-title..work/input/.Do not silently switch ASR providers, use local Whisper, add fallback model paths, install system tools, or reveal secrets.
*.maas.aliyuncs.com HTTPS endpoints; Agent mode processes transcript sections only through the already-selected host Agent model and adds no caller-supplied endpoint. Do not accept arbitrary upload or model endpoints.For standalone use, run python scripts/preflight.py and send stdout verbatim. Do not paraphrase, reorder, add options, or ask whether the user wants Simplified Chinese, Traditional Chinese, or bilingual subtitles. Simplified Chinese is the default target; bilingual ASS/SRT is the fixed output structure. In video-flow, reuse video-download/scripts/preflight.py --mode combined only for a remote-URL route. A local video or local media-project route skips video-download, runs this Skill's questionnaire, and never asks about download quality.
Run commands from this skill folder. On a new device or unverified environment, run the local-only check:
python scripts/check_env.py
Confirm only these user-facing inputs unless already clear:
--language en).outputs/; confirm any different path. When the media came from the download Skill, use its media project folder so final ASS/SRT and hidden .work/ artifacts stay together. Hidden audio and source subtitles are temporary and are removed only after successful export.qwen-mt-plus for stable, cost-effective API translation. When the user explicitly chooses Codex / Agent translation, add --translation-provider agent; this consumes the current Agent's model allowance and requires no additional translation API Key.https://www.okfile.com and its temporary URL is sent to Alibaba Fun-ASR. With qwen-mt-plus, subtitle text is also sent to Alibaba; with Agent mode, transcript text is handled by the current Agent model service. Do not proceed without an affirmative answer.Do not ask ordinary users to choose ASR, segment-generation, or orchestration models.
Output naming is automatic unless more than one plausible source video/title remains:
[platform-id], extension, and subtitle/release suffix. Never use 原版视频, 原视频, 视频, source video, or another generic placeholder.--localized-title. Do not ask the user to translate an unambiguous title or choose a filename.output_naming.json; resumed and final export commands must match it exactly.3 through 6 are immediate Agent work gates, not reasons to end the user task. Complete the applicable gate and rerun with the same run ID and translation provider in the same turn.Start a normal run with:
python scripts/video_to_subtitles.py "/absolute/path/to/video.mp4" \
--localized-title "<按领域术语翻译的中文净标题>" \
--language en \
--confirm-external-processing
The public default above uses qwen-mt-plus. When the user explicitly selects current Codex / Agent translation, add:
--translation-provider agent
Add --outputs-dir "<project-path>" after the user confirms the media project folder. The default working directory becomes <project-path>/.work/, keeping intermediate files out of the Skill source directory. For screen-recording guidance, read screen context rules before generating screenshots.
In a combined workflow, the hidden .work/input/ audio and source subtitle are discovered automatically. Use --source-subtitle "/absolute/path/reference.srt" only when the reference is outside the standard project layout. Use --keep-workflow-inputs only for explicit debugging; normal successful delivery removes those temporary inputs.
The wrapper uses hash-bound Agent gates: exit 3 is whole-source analysis; Agent translation adds exit 4; exit 5 is whole-document semantic translation review; exit 6 is final whole-document QC. The default qwen path skips exit 4. Follow each generated WORKFLOW.md, complete its receipt, and rerun the same command with the same --run-id and --translation-provider without ending the user task.
For other failures, use workflow_status.json, final_qa_report.md, final_qa_prompt.txt, and python scripts/check_env.py --json. Repair the affected semantic-review section files automatically before asking the user; ask only after two failed repair attempts or when domain judgment is necessary.
Do not call export_subtitles.py as a shortcut or export while any workflow step is waiting/running. Delivery requires current, hash-matching source-analysis.validated.json, Agent-translation validation when applicable, semantic-review.validated.json, final_qa.validated.json, final_qa_report.md with Blockers: 0, and final-qc.validated.json. Final QC must spot-check the opening 30 cues, a middle/numeric passage, core domain terms, direction/entry/add/cover logic, names/tickers/amounts, sponsorship, and user-flagged timestamps, using a reasoned not_applicable only when appropriate.
In every SRT cue, place Chinese and source text on separate physical lines; never write literal /n, \\n, \\N, <br>, or ASS tags into SRT text. Deliver only <中文净标题>.<中X双语字幕>.ass and the matching .srt; the title must contain the real localized source title with no leading date, trailing platform ID, extension, release suffix, or generic placeholder. BCC is not an output of this public Skill. After success, report the ASS path, SRT path, elapsed time, models used, QA blocker/warning counts, and any focused spot-check recommendation.
The repository-level product guide is outside the installable skill package. Do not treat product documentation as the execution contract.