mc-agent
無料Front the studio as Manny the Manticore. Use when the user says "Manny", "Manticore", or "talk to Manny".
日本語の概要は準備中です。原文の説明を表示しています。
Farm narration, music beds, and SFX locally. Use when another skill needs sound, or when the user says "add narration", "music bed", or "sound effect".
インストール方法を見るインストールする前に、エージェントに与えられる指示の中身を確認できます。
mc-assets farms pictures; this skill farms sound. A caller (another skill, or the creator directly) says what sound is needed and where it lands; you resolve the lane, run the engine, and hand back files with provenance. It is a service skill: it owns no stage, stops at no gate, and writes no project state. The caller mixes what you hand back without hearing it first, so the bar is honesty about what came back.
Read {skill-root}/references/audio-lanes.md before farming anything; the ladder, the limits, and the honesty rules there are binding.
{video-path}, the current video project at {projects-path}/<slug>/.{skill-root} → this skill's installed directory; files in it always carry it ({skill-root}/references/audio-lanes.md).{project-root} → the project working directory.uv run {project-root}/_bmad/scripts/resolve_config.py --project-root {project-root} --key modules.manticore. Empty means mc-setup has not run: stop and route the creator there.paths values against {project-root}. From [audio] take the lane values (tts-provider, music-provider, sfx-provider, song-provider) and workspace; the engine workspace is {engines-path}/{audio.workspace}.The implemented 1.0 lanes are the local defaults: kokoro-local (tts and podcast), musicgen-local (music), audioldm2-local (sfx). A paid or planned value (gemini-tts, elevenlabs-*, stable-audio-open, ace-step-local) means the creator opted into a lane that has not landed: say so plainly and stop. Never substitute a paid lane the creator did not choose, and never pretend an unvalidated lane works. An empty song-provider is the shipped state, not a misconfiguration: full songs with vocals have no validated local lane yet.
uv run {skill-root}/scripts/ensure_workspace.py --workspace <resolved workspace> --check. Not ready: tell the creator what a bootstrap downloads (venv wheels of several GB; on Windows with an NVIDIA GPU, torch installs CUDA wheels from the PyTorch cu126 index, adding roughly 2.5 to 3 GB more; ~340 MB of Kokoro models now, ~5 GB of Hugging Face cache on the first music/sfx run), relay the torch field of the script's --dry-run JSON verbatim so they know which wheel source this machine will use, get their go-ahead, then run it without --check. An existing validated workspace is used as-is, a lab the creator built by hand included: never rebuilt, never duplicated.
One call per asset:
uv run {skill-root}/scripts/farm_audio.py --kind tts|podcast|music|sfx --provider <the [audio] lane value> --workspace <resolved workspace> --out-dir <where the caller wants it> [--name <basename>] plus the kind's arguments (--text/--voice/--speed, --script lines.json, --prompt/--seconds/--seed). The script appends provenance to <out-dir>/manifest.json (same row shape as mc-assets; cost is null on local lanes).
Podcast dialogue takes a script JSON in the shape the reference gives, with its realism knobs applied (speed variation, gaps, backchannels) rather than uniform lines.
First-run model downloads are long: run them in the background with proactive progress reports.
Listen to or inspect every output before presenting it: duration matches the request, nothing silent or truncated, dialogue lines in order. Deliver with the caveats from the reference that apply: SFX are 16 kHz (fine under a mix, thin exposed solo), music is instrumental only, TTS voices are stock (no cloning, no "your voice" claims), crosstalk is simulated. Report where every file landed.
[audio] in the studio config; no paid or metered lane ran without the creator's explicit configuration, and no planned lane was presented as working.まだレビューはありません。使ってみた感想をお寄せください。
概要と使いどころ
Front the studio as Manny the Manticore. Use when the user says "Manny", "Manticore", or "talk to Manny".
日本語の概要は準備中です。原文の説明を表示しています。
Farm the stills and b-roll the beats need. Use at the assets stage, or when the user says "farm the assets", "find the images", or "get the b-roll".
日本語の概要は準備中です。原文の説明を表示しています。
Riff visuals, then build the graphics beat table. Use at the beats stage after gate 2, or when the user says "plan the graphics", "beat table", or "what visuals go here".
日本語の概要は準備中です。原文の説明を表示しています。
Capture the creator's idea in their exact words. Use at the braindump stage, or when the user says "braindump", "let me talk this through", or "here is my idea".
日本語の概要は準備中です。原文の説明を表示しています。
Cut raw takes into an approved, rendered edit. Use at the cut stage with recordings in raw/, or when the user says "cut the takes", "make the cutplan", "render the preview", or "render the final".
日本語の概要は準備中です。原文の説明を表示しています。
Render the beat table into alpha overlays. Use at the graphics stage after gate 3, or when the user says "build the graphics" or "render the overlays".
日本語の概要は準備中です。原文の説明を表示しています。