AI剪口播
無料口播视频转录和口误识别。生成审查稿和删除任务清单。触发词:剪口播、处理视频、识别口误
日本語の概要は準備中です。原文の説明を表示しています。
Generate cloned narration for the content workspace. Use local IndexTTS2 with the canonical lossless Pluviobyte reference for every local video, voice audition, digital-human narration, and subtitle-ready audio render. Preserve MiniMax only for public-facing articles or tutorials that teach users to connect through the relay service; never use MiniMax for local video production.
インストールする前に、エージェントに与えられる指示の中身を確認できます。
Route every request into exactly one lane.
.env.pluvio-indextts2-calm-v1 and the canonical lossless WAV in
automation/config/tts-routing.json. Never use an MP3 as the speaker
reference and never replace the canonical reference with the latest output;
that would accumulate cloning drift.default_delivery.playback_speed declared in
automation/config/tts-routing.json to every audition and final WAV. The
production default is 1.12x, implemented with FFmpeg atempo so pitch is
preserved. Record the effective multiplier in voice_manifest.json.say, browser speech, MiniMax, or another
generic/cloud voice. If IndexTTS2, its model weights, Apple MPS/CPU, or the
canonical reference is unavailable, repair the local lane or stop.voice_manifest.json with provider,
voice id, model, canonical reference path and SHA-256, segment contract,
output path, and used_fallback=false.ra-audio-to-subtitles against the exact concatenated lossless narration and
use its word timestamps for final captions.references/pronunciation-lexicon.json. A user-approved pronunciation in
that file overrides generic phonetic heuristics. Keep the display spelling
separate from TTS text when needed, and stop rather than render a forbidden
spelling.Read the machine-readable route from:
automation/config/tts-routing.json
Generate a narration from a JSONL segment contract with:
python3 .claude/skills/tts-skill/scripts/generate_indextts2_narration.py \
--batch-file <segments.jsonl> \
--output <final-narration.wav> \
--manifest <voice_manifest.json>
Each non-empty JSONL line must follow the official IndexTTS2 batch contract and
contain text; it may also contain emotion_vector, emotion_weight, and
silence_after_ms. Default to a calm delivery. Add stronger emotion only when
the user explicitly asks for it.
Run --dry-run before a new contract or after changing the local model setup.
The helper verifies that the canonical speaker reference is a lossless PCM WAV,
invokes the official local IndexTTS2 CLI, normalizes the result, and writes the
manifest. Keep the raw WAV beside the normalized WAV for auditability.
For mixed Chinese/English scripts, inspect model names, brands, abbreviations,
and numbers before the full render. Load
references/pronunciation-lexicon.json first. Entries there are approved
production contracts: use their tts value exactly and reject every
forbidden_tts form. In particular, keep Codex as the raw English token;
never transliterate it as “扣代克斯”, “扣戴克斯”, or “扣德克斯”. For a term
that is not yet in the lexicon, audition alternatives while keeping canonical
spelling in the display script and final captions. After the user chooses,
record that decision in the lexicon before the full render.
MiniMax is retained only as the public solution used in articles, tutorials,
course examples, and relay-station onboarding. It is not a local production
provider. When this lane is requested, read
references/minimax-relay-article.md and use assets/minimax_tts.py only for
the article/demo example.
Before using narration in a final video:
provider is IndexTTS2/indextts2-localvoice_id is pluvio-indextts2-calm-v1playback_speed is 1.12 unless the production contract explicitly
overrides the workspace defaultused_fallback=false and no segment provider is MiniMaxvoice_manifest.json and
the approved mixed-language terms were usedcaption-qc.json PASSまだレビューはありません。使ってみた感想をお寄せください。
概要と使いどころ
口播视频转录和口误识别。生成审查稿和删除任务清单。触发词:剪口播、处理视频、识别口误
日本語の概要は準備中です。原文の説明を表示しています。
将法天象地、同源巨大法相或角色力量显现的视觉参考制作成连续关键帧,再使用用户选择或可用的图生视频工具分段生成、检查和合成动作视频,不绑定成片平台。适用于本体与法相同框、领域展开、凝实与同步出招;不负责整部小说改编或普通图片轮播。
日本語の概要は準備中です。原文の説明を表示しています。
口播基础素材包生成。转录口播、识别口误、生成审核页;用户确认后剪出新视频,Agent 再基于剪后视频重新转写、AI 校对字幕,输出后续口播成片可用的 source_cut.mp4 和 subtitles.srt。触发词:剪口播、处理口播素材、准备口播素材、识别口误、基础素材包
日本語の概要は準備中です。原文の説明を表示しています。
口播视频成片 Skill。把文章/口播稿/SRT、剪后视频和 HTML/图片素材串成分镜稿、时间线预览和最终 MP4;成片比例和动画风格从用户配置读取,动画默认使用小黑风格。触发词:口播成片、做分镜稿、时间线预览、合成口播视频、导出竖屏MP4
日本語の概要は準備中です。原文の説明を表示しています。
自进化 skills。记录用户反馈,更新方法论和规则。触发词:更新规则、记录反馈、改进skill
日本語の概要は準備中です。原文の説明を表示しています。
dontbesilent 商业工具箱主入口。双模式:任务前路由(你的问题该用哪个 skill)+ 任务后导航(刚做完诊断,下一步该干什么)。 触发方式:/dbs、/商业、「帮我看看」、「下一步怎么走」 Main entry point for dontbesilent business toolkit. Dual mode: pre-task routing + post-task navigation. Trigger: /dbs, "help me with my business", "what's next"
日本語の概要は準備中です。原文の説明を表示しています。