Web・iOS・Androidの画面を、読み上げやキーボード操作に対応させ、ラベル、配色、操作対象の大きさなどをWCAG 2.2に沿って設計・点検するスキル。
- アイコンボタンの説明を付けたいとき
- キーボード操作とモーダルの点検
- コントラストや操作対象の大きさの確認
撮影済みの動画や画面収録から見どころを選び、カット・構成整理・字幕や音声の追加を進めます。長い録画の短編化からSNS向けの画角調整、編集ソフトでの仕上げまで扱うスキルです。
AI-assisted video editing workflows for cutting, structuring, and augmenting real footage. Covers the full pipeline from raw capture through FFmpeg, Remotion, ElevenLabs, fal.ai, and final polish in Descript or CapCut. Use when the user wants to edit video, cut footage, create vlogs, or build video content.
撮影済みの動画や画面収録を整理し、見どころを選んで一本の動画にまとめる作業を支援します。文字起こしから不要な間や重複を探し、残す区間と順番を決めます。FFmpegで切り出しや結合、音量調整を行い、Remotionで文字や図表を重ねます。必要に応じてナレーション、音楽、補足画像を追加し、SNS向けの縦長・正方形への画角調整も扱います。
長時間の録画から短い見どころ動画を作りたいときや、Vlog、チュートリアル、製品デモを組み立てたいときに向いています。同じ形式の演出を複数の動画で使い回す作業にも役立ちます。
実写素材の編集が中心で、動画全体をプロンプトから生成する用途ではありません。工程に応じてFFmpeg、Remotionなどの外部ツールを使い、ElevenLabsのAPI利用にはAPIキーが必要です。最後のテンポ、字幕修正、色調、音声バランスはDescriptやCapCutで人が確認・調整します。
この紹介文は、公開されている SKILL.md をもとに AI(Claude Haiku)が作成しました。正確な仕様は下の原文を確認してください。
インストールする前に、エージェントに与えられる指示の中身を確認できます。
AI-assisted editing for real footage. Not generation from prompts. Editing existing video fast.
AI video editing is useful when you stop asking it to create the whole video and start using it to compress, structure, and augment real footage. The value is not generation. The value is compression.
For measured reference-driven work, chain taste-distillation into
taste-application, then return here for the editor and final-output review.
The standalone taste skills can use existing footage; generation is optional.
Before live editor or DAW changes, save a versioned project checkpoint and verify the file exists. Save and verify another checkpoint after the changes. An API readback proves the current in-memory state, not that it was saved. Keep rendered media, editable projects, and creative approval as separate states in the handoff.
For MIDI-driven audio, check pitches against the receiving rack's note mapping and audition the result; successful clip creation can still produce silence. For reconstructed projects, validate through native load and save, sort events in timeline order, verify sample links and mute states, then check and audition the exact exported audio for unintended silence. XML parsing alone does not prove that the DAW accepted every clip or produced audible output. Check a bridge's capability handshake before invoking newer commands. Do not enable upload or training-data telemetry as a side effect of a creative task; use a supported local control path when consent or capability is absent.
Screen Studio / raw footage
→ Claude / Codex
→ FFmpeg
→ Remotion
→ ElevenLabs / fal.ai
→ Descript or CapCut
Each layer has a specific job. Do not skip layers. Do not try to make one tool do everything.
Collect the source material:
videodb skill)Output: raw files ready for organization.
Use Claude Code or Codex to:
Example prompt:
"Here's the transcript of a 4-hour recording. Identify the 8 strongest segments
for a 24-minute vlog. Give me FFmpeg cut commands for each segment."
This layer is about structure, not final creative taste.
FFmpeg handles the boring but critical work: splitting, trimming, concatenating, and preprocessing.
ffmpeg -i raw.mp4 -ss 00:12:30 -to 00:15:45 -c copy segment_01.mp4
#!/bin/bash
# cuts.txt: start,end,label
while IFS=, read -r start end label; do
ffmpeg -i raw.mp4 -ss "$start" -to "$end" -c copy "segments/${label}.mp4"
done < cuts.txt
# Create file list
for f in segments/*.mp4; do echo "file '$f'"; done > concat.txt
ffmpeg -f concat -safe 0 -i concat.txt -c copy assembled.mp4
ffmpeg -i raw.mp4 -vf "scale=960:-2" -c:v libx264 -preset ultrafast -crf 28 proxy.mp4
ffmpeg -i raw.mp4 -vn -acodec pcm_s16le -ar 16000 audio.wav
ffmpeg -i segment.mp4 -af loudnorm=I=-16:TP=-1.5:LRA=11 -c:v copy normalized.mp4
Remotion turns editing problems into composable code. Use it for things that traditional editors make painful:
import { AbsoluteFill, Sequence, Video, useCurrentFrame } from "remotion";
export const VlogComposition: React.FC = () => {
const frame = useCurrentFrame();
return (
<AbsoluteFill>
{/* Main footage */}
<Sequence from={0} durationInFrames={300}>
<Video src="/segments/intro.mp4" />
</Sequence>
{/* Title overlay */}
<Sequence from={30} durationInFrames={90}>
<AbsoluteFill style={{
justifyContent: "center",
alignItems: "center",
}}>
<h1 style={{
fontSize: 72,
color: "white",
textShadow: "2px 2px 8px rgba(0,0,0,0.8)",
}}>
The AI Editing Stack
</h1>
</AbsoluteFill>
</Sequence>
{/* Next segment */}
<Sequence from={300} durationInFrames={450}>
<Video src="/segments/demo.mp4" />
</Sequence>
</AbsoluteFill>
);
};
npx remotion render src/index.ts VlogComposition output.mp4
See the Remotion docs for detailed patterns and API reference.
Generate only what you need. Do not generate the whole video.
import os
import requests
resp = requests.post(
f"https://api.elevenlabs.io/v1/text-to-speech/{voice_id}",
headers={
"xi-api-key": os.environ["ELEVENLABS_API_KEY"],
"Content-Type": "application/json"
},
json={
"text": "Your narration text here",
"model_id": "eleven_turbo_v2_5",
"voice_settings": {"stability": 0.5, "similarity_boost": 0.75}
}
)
with open("voiceover.mp3", "wb") as f:
f.write(resp.content)
Use the fal-ai-media skill for:
Use for insert shots, thumbnails, or b-roll that doesn't exist:
generate(app_id: "fal-ai/nano-banana-pro", input_data: {
"prompt": "professional thumbnail for tech vlog, dark background, code on screen",
"image_size": "landscape_16_9"
})
If VideoDB is configured:
voiceover = coll.generate_voice(text="Narration here", voice="alloy")
music = coll.generate_music(prompt="lo-fi background for coding vlog", duration=120)
sfx = coll.generate_sound_effect(prompt="subtle whoosh transition")
The last layer is human. Use a traditional editor for:
This is where taste lives. AI clears the repetitive work. You make the final calls.
Different platforms need different aspect ratios:
| Platform | Aspect Ratio | Resolution |
|---|---|---|
| YouTube | 16:9 | 1920x1080 |
| TikTok / Reels | 9:16 | 1080x1920 |
| Instagram Feed | 1:1 | 1080x1080 |
| X / Twitter | 16:9 or 1:1 | 1280x720 or 720x720 |
# 16:9 to 9:16 (center crop)
ffmpeg -i input.mp4 -vf "crop=ih*9/16:ih,scale=1080:1920" vertical.mp4
# 16:9 to 1:1 (center crop)
ffmpeg -i input.mp4 -vf "crop=ih:ih,scale=1080:1080" square.mp4
from videodb import ReframeMode
# Smart reframe (AI-guided subject tracking)
reframed = video.reframe(start=0, end=60, target="vertical", mode=ReframeMode.smart)
# Detect scene changes (threshold 0.3 = moderate sensitivity)
ffmpeg -i input.mp4 -vf "select='gt(scene,0.3)',showinfo" -vsync vfr -f null - 2>&1 | grep showinfo
# Find silent segments (useful for cutting dead air)
ffmpeg -i input.mp4 -af silencedetect=noise=-30dB:d=2 -f null - 2>&1 | grep silence
Use Claude to analyze transcript + scene timestamps:
"Given this transcript with timestamps and these scene change points,
identify the 5 most engaging 30-second clips for social media."
| Tool | Strength | Weakness |
|---|---|---|
| Claude / Codex | Organization, planning, code generation | Not the creative taste layer |
| FFmpeg | Deterministic cuts, batch processing, format conversion | No visual editing UI |
| Remotion | Programmable overlays, composable scenes, reusable templates | Learning curve for non-devs |
| Screen Studio | Polished screen recordings immediately | Only screen capture |
| ElevenLabs | Voice, narration, music, SFX | Not the center of the workflow |
| Descript / CapCut | Final pacing, captions, polish | Manual, not automatable |
ITO Production v1 provides restrained highlight bloom, opposing RGB spatial offsets and a luminance/edge halo. The exact files passed prior native import, save/reopen and short motion-render checks after two-source visual review. These are starting values requiring shot-specific review; the halo does not detect or track subjects.
ITO V28 contains preserved, native-verified Fusion graph snippets and an idempotent Lua installer. These are technical compatibility examples, not recommended production defaults: their documented visual limitations require tuning and taste review before use. See the bundle provenance for the scope of prior import and render checks.
fal-ai-media — AI image, video, and audio generationvideodb — Server-side video processing, indexing, and streamingcontent-engine — Platform-native content distributionまだレビューはありません。使ってみた感想をお寄せください。
概要と使いどころ
Web・iOS・Androidの画面を、読み上げやキーボード操作に対応させ、ラベル、配色、操作対象の大きさなどをWCAG 2.2に沿って設計・点検するスキル。
AIエージェントの不調を、指示・記憶・ツール実行・画面表示など12の層から調べるスキル。コードやログを根拠に原因を整理し、重要度順の指摘と修正案をまとめます。
実際の開発課題で複数のコーディングエージェントを比較するスキル。成功率、取得可能なAPI費用、所要時間、繰り返し実行の安定性を測り、選定や更新後の評価に使えます。
AIエージェントが使うツールの種類や入出力、エラーからの復帰手順を設計・見直します。文脈の情報量も整理し、作業完了率や再試行回数で改善を評価します。
AIエージェントの失敗や同じ操作の繰り返しを記録し、原因の切り分け、小さな復旧操作、結果の報告まで進める手順を示して、根拠のある再試行につなげるスキル。
AIエージェントが失敗や同じ操作を繰り返す原因を、エラーと実行状況から整理します。小さな復旧操作を試し、結果と根拠を引き継げる報告にまとめるスキルです。