本文へ移動
cccskills
無料GitHub で公開日本語紹介

video-editing

撮影済みの動画や画面収録から見どころを選び、カット・構成整理・字幕や音声の追加を進めます。長い録画の短編化からSNS向けの画角調整、編集ソフトでの仕上げまで扱うスキルです。

原文原文の説明を見る

AI-assisted video editing workflows for cutting, structuring, and augmenting real footage. Covers the full pipeline from raw capture through FFmpeg, Remotion, ElevenLabs, fal.ai, and final polish in Descript or CapCut. Use when the user wants to edit video, cut footage, create vlogs, or build video content.

インストール方法を見る

こんなときに便利

  • 長い録画から見どころを切り出したいとき
  • 撮影素材からVlogを作りたいとき
  • 製品デモに文字やナレーションを加えたいとき
  • SNS向けに動画を縦長にしたいとき

日本語での紹介

できること

撮影済みの動画や画面収録を整理し、見どころを選んで一本の動画にまとめる作業を支援します。文字起こしから不要な間や重複を探し、残す区間と順番を決めます。FFmpegで切り出しや結合、音量調整を行い、Remotionで文字や図表を重ねます。必要に応じてナレーション、音楽、補足画像を追加し、SNS向けの縦長・正方形への画角調整も扱います。

こんなときに便利

長時間の録画から短い見どころ動画を作りたいときや、Vlog、チュートリアル、製品デモを組み立てたいときに向いています。同じ形式の演出を複数の動画で使い回す作業にも役立ちます。

使い方の例

  • 「この4時間の録画の文字起こしから、24分のVlogに残す8区間を選んで」
  • 「この動画を縦長に切り抜き、タイトルとナレーションを追加して」

注意点

実写素材の編集が中心で、動画全体をプロンプトから生成する用途ではありません。工程に応じてFFmpeg、Remotionなどの外部ツールを使い、ElevenLabsのAPI利用にはAPIキーが必要です。最後のテンポ、字幕修正、色調、音声バランスはDescriptやCapCutで人が確認・調整します。

この紹介文は、公開されている SKILL.md をもとに AI(Claude Haiku)が作成しました。正確な仕様は下の原文を確認してください。

含まれるファイル(13)

  • SKILL.md11.5 KB
  • assets/fusion/ito-production-v1/install_ito_production_v1.lua2.1 KB
  • assets/fusion/ito-production-v1/ITO_PROD_HighlightBloom.setting344 B
  • assets/fusion/ito-production-v1/ITO_PROD_LumaHalo.setting875 B
  • assets/fusion/ito-production-v1/ITO_PROD_RGBFringe.setting952 B
  • assets/fusion/ito-production-v1/provenance.json2.3 KB
  • assets/fusion/ito-production-v1/README.md3.1 KB
  • assets/fusion/ito-v28/install_ito_v28.lua2.1 KB
  • assets/fusion/ito-v28/ITO_V28_FlashEtherealBloom.setting572 B
  • assets/fusion/ito-v28/ITO_V28_RGBDisplacement.setting592 B
  • assets/fusion/ito-v28/ITO_V28_SubjectHalo.setting582 B
  • assets/fusion/ito-v28/provenance.json1.8 KB
  • assets/fusion/ito-v28/README.md3.1 KB

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

Video Editing

AI-assisted editing for real footage. Not generation from prompts. Editing existing video fast.

When to Activate

  • User wants to edit, cut, or structure video footage
  • Turning long recordings into short-form content
  • Building vlogs, tutorials, or demo videos from raw capture
  • Adding overlays, subtitles, music, or voiceover to existing video
  • Reframing video for different platforms (YouTube, TikTok, Instagram)
  • User says "edit video", "cut this footage", "make a vlog", or "video workflow"

Core Thesis

AI video editing is useful when you stop asking it to create the whole video and start using it to compress, structure, and augment real footage. The value is not generation. The value is compression.

The Pipeline

For measured reference-driven work, chain taste-distillation into taste-application, then return here for the editor and final-output review. The standalone taste skills can use existing footage; generation is optional.

Before live editor or DAW changes, save a versioned project checkpoint and verify the file exists. Save and verify another checkpoint after the changes. An API readback proves the current in-memory state, not that it was saved. Keep rendered media, editable projects, and creative approval as separate states in the handoff.

For MIDI-driven audio, check pitches against the receiving rack's note mapping and audition the result; successful clip creation can still produce silence. For reconstructed projects, validate through native load and save, sort events in timeline order, verify sample links and mute states, then check and audition the exact exported audio for unintended silence. XML parsing alone does not prove that the DAW accepted every clip or produced audible output. Check a bridge's capability handshake before invoking newer commands. Do not enable upload or training-data telemetry as a side effect of a creative task; use a supported local control path when consent or capability is absent.

Screen Studio / raw footage
  → Claude / Codex
  → FFmpeg
  → Remotion
  → ElevenLabs / fal.ai
  → Descript or CapCut

Each layer has a specific job. Do not skip layers. Do not try to make one tool do everything.

Layer 1: Capture (Screen Studio / Raw Footage)

Collect the source material:

  • Screen Studio: polished screen recordings for app demos, coding sessions, browser workflows
  • Raw camera footage: vlog footage, interviews, event recordings
  • Desktop capture via VideoDB: session recording with real-time context (see videodb skill)

Output: raw files ready for organization.

Layer 2: Organization (Claude / Codex)

Use Claude Code or Codex to:

  • Transcribe and label: generate transcript, identify topics and themes
  • Plan structure: decide what stays, what gets cut, what order works
  • Identify dead sections: find pauses, tangents, repeated takes
  • Generate edit decision list: timestamps for cuts, segments to keep
  • Scaffold FFmpeg and Remotion code: generate the commands and compositions
Example prompt:
"Here's the transcript of a 4-hour recording. Identify the 8 strongest segments
for a 24-minute vlog. Give me FFmpeg cut commands for each segment."

This layer is about structure, not final creative taste.

Layer 3: Deterministic Cuts (FFmpeg)

FFmpeg handles the boring but critical work: splitting, trimming, concatenating, and preprocessing.

Extract segment by timestamp

ffmpeg -i raw.mp4 -ss 00:12:30 -to 00:15:45 -c copy segment_01.mp4

Batch cut from edit decision list

#!/bin/bash
# cuts.txt: start,end,label
while IFS=, read -r start end label; do
  ffmpeg -i raw.mp4 -ss "$start" -to "$end" -c copy "segments/${label}.mp4"
done < cuts.txt

Concatenate segments

# Create file list
for f in segments/*.mp4; do echo "file '$f'"; done > concat.txt
ffmpeg -f concat -safe 0 -i concat.txt -c copy assembled.mp4

Create proxy for faster editing

ffmpeg -i raw.mp4 -vf "scale=960:-2" -c:v libx264 -preset ultrafast -crf 28 proxy.mp4

Extract audio for transcription

ffmpeg -i raw.mp4 -vn -acodec pcm_s16le -ar 16000 audio.wav

Normalize audio levels

ffmpeg -i segment.mp4 -af loudnorm=I=-16:TP=-1.5:LRA=11 -c:v copy normalized.mp4

Layer 4: Programmable Composition (Remotion)

Remotion turns editing problems into composable code. Use it for things that traditional editors make painful:

When to use Remotion

  • Overlays: text, images, branding, lower thirds
  • Data visualizations: charts, stats, animated numbers
  • Motion graphics: transitions, explainer animations
  • Composable scenes: reusable templates across videos
  • Product demos: annotated screenshots, UI highlights

Basic Remotion composition

import { AbsoluteFill, Sequence, Video, useCurrentFrame } from "remotion";

export const VlogComposition: React.FC = () => {
  const frame = useCurrentFrame();

  return (
    <AbsoluteFill>
      {/* Main footage */}
      <Sequence from={0} durationInFrames={300}>
        <Video src="/segments/intro.mp4" />
      </Sequence>

      {/* Title overlay */}
      <Sequence from={30} durationInFrames={90}>
        <AbsoluteFill style={{
          justifyContent: "center",
          alignItems: "center",
        }}>
          <h1 style={{
            fontSize: 72,
            color: "white",
            textShadow: "2px 2px 8px rgba(0,0,0,0.8)",
          }}>
            The AI Editing Stack
          </h1>
        </AbsoluteFill>
      </Sequence>

      {/* Next segment */}
      <Sequence from={300} durationInFrames={450}>
        <Video src="/segments/demo.mp4" />
      </Sequence>
    </AbsoluteFill>
  );
};

Render output

npx remotion render src/index.ts VlogComposition output.mp4

See the Remotion docs for detailed patterns and API reference.

Layer 5: Generated Assets (ElevenLabs / fal.ai)

Generate only what you need. Do not generate the whole video.

Voiceover with ElevenLabs

import os
import requests

resp = requests.post(
    f"https://api.elevenlabs.io/v1/text-to-speech/{voice_id}",
    headers={
        "xi-api-key": os.environ["ELEVENLABS_API_KEY"],
        "Content-Type": "application/json"
    },
    json={
        "text": "Your narration text here",
        "model_id": "eleven_turbo_v2_5",
        "voice_settings": {"stability": 0.5, "similarity_boost": 0.75}
    }
)
with open("voiceover.mp3", "wb") as f:
    f.write(resp.content)

Music and SFX with fal.ai

Use the fal-ai-media skill for:

  • Background music generation
  • Sound effects (ThinkSound model for video-to-audio)
  • Transition sounds

Generated visuals with fal.ai

Use for insert shots, thumbnails, or b-roll that doesn't exist:

generate(app_id: "fal-ai/nano-banana-pro", input_data: {
  "prompt": "professional thumbnail for tech vlog, dark background, code on screen",
  "image_size": "landscape_16_9"
})

VideoDB generative audio

If VideoDB is configured:

voiceover = coll.generate_voice(text="Narration here", voice="alloy")
music = coll.generate_music(prompt="lo-fi background for coding vlog", duration=120)
sfx = coll.generate_sound_effect(prompt="subtle whoosh transition")

Layer 6: Final Polish (Descript / CapCut)

The last layer is human. Use a traditional editor for:

  • Pacing: adjust cuts that feel too fast or slow
  • Captions: auto-generated, then manually cleaned
  • Color grading: basic correction and mood
  • Final audio mix: balance voice, music, and SFX levels
  • Export: platform-specific formats and quality settings

This is where taste lives. AI clears the repetitive work. You make the final calls.

Social Media Reframing

Different platforms need different aspect ratios:

PlatformAspect RatioResolution
YouTube16:91920x1080
TikTok / Reels9:161080x1920
Instagram Feed1:11080x1080
X / Twitter16:9 or 1:11280x720 or 720x720

Reframe with FFmpeg

# 16:9 to 9:16 (center crop)
ffmpeg -i input.mp4 -vf "crop=ih*9/16:ih,scale=1080:1920" vertical.mp4

# 16:9 to 1:1 (center crop)
ffmpeg -i input.mp4 -vf "crop=ih:ih,scale=1080:1080" square.mp4

Reframe with VideoDB

from videodb import ReframeMode

# Smart reframe (AI-guided subject tracking)
reframed = video.reframe(start=0, end=60, target="vertical", mode=ReframeMode.smart)

Scene Detection and Auto-Cut

FFmpeg scene detection

# Detect scene changes (threshold 0.3 = moderate sensitivity)
ffmpeg -i input.mp4 -vf "select='gt(scene,0.3)',showinfo" -vsync vfr -f null - 2>&1 | grep showinfo

Silence detection for auto-cut

# Find silent segments (useful for cutting dead air)
ffmpeg -i input.mp4 -af silencedetect=noise=-30dB:d=2 -f null - 2>&1 | grep silence

Highlight extraction

Use Claude to analyze transcript + scene timestamps:

"Given this transcript with timestamps and these scene change points,
identify the 5 most engaging 30-second clips for social media."

What Each Tool Does Best

ToolStrengthWeakness
Claude / CodexOrganization, planning, code generationNot the creative taste layer
FFmpegDeterministic cuts, batch processing, format conversionNo visual editing UI
RemotionProgrammable overlays, composable scenes, reusable templatesLearning curve for non-devs
Screen StudioPolished screen recordings immediatelyOnly screen capture
ElevenLabsVoice, narration, music, SFXNot the center of the workflow
Descript / CapCutFinal pacing, captions, polishManual, not automatable

Key Principles

  1. Edit, don't generate. This workflow is for cutting real footage, not creating from prompts.
  2. Structure before style. Get the story right in Layer 2 before touching anything visual.
  3. FFmpeg is the backbone. Boring but critical. Where long footage becomes manageable.
  4. Remotion for repeatability. If you'll do it more than once, make it a Remotion component.
  5. Generate selectively. Only use AI generation for assets that don't exist, not for everything.
  6. Taste is the last layer. AI clears repetitive work. You make the final creative calls.

Native Fusion Presets

ITO Production v1 provides restrained highlight bloom, opposing RGB spatial offsets and a luminance/edge halo. The exact files passed prior native import, save/reopen and short motion-render checks after two-source visual review. These are starting values requiring shot-specific review; the halo does not detect or track subjects.

ITO V28 contains preserved, native-verified Fusion graph snippets and an idempotent Lua installer. These are technical compatibility examples, not recommended production defaults: their documented visual limitations require tuning and taste review before use. See the bundle provenance for the scope of prior import and render checks.

Related Skills

  • fal-ai-media — AI image, video, and audio generation
  • videodb — Server-side video processing, indexing, and streaming
  • content-engine — Platform-native content distribution

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

同じリポジトリのスキル

概要と使いどころ

accessibility

無料日本語概要

Web・iOS・Androidの画面を、読み上げやキーボード操作に対応させ、ラベル、配色、操作対象の大きさなどをWCAG 2.2に沿って設計・点検するスキル。

  • アイコンボタンの説明を付けたいとき
  • キーボード操作とモーダルの点検
  • コントラストや操作対象の大きさの確認
affaan-m/ECC27.7万2026年10月12日 更新

agent-architecture-audit

無料日本語概要

AIエージェントの不調を、指示・記憶・ツール実行・画面表示など12の層から調べるスキル。コードやログを根拠に原因を整理し、重要度順の指摘と修正案をまとめます。

  • アプリ内だけで起きる不調を調べたいとき
  • 過去の会話が混ざる原因を調査
  • ツールの未実行や実行の誤報を確認
affaan-m/ECC27.7万2026年10月12日 更新

agent-eval

無料日本語概要

実際の開発課題で複数のコーディングエージェントを比較するスキル。成功率、取得可能なAPI費用、所要時間、繰り返し実行の安定性を測り、選定や更新後の評価に使えます。

  • 実際の開発課題でエージェントを比較
  • 新しいツールやモデルの導入前評価
  • エージェント更新後の性能確認
affaan-m/ECC27.7万2026年10月12日 更新

agent-harness-construction

無料日本語概要

AIエージェントが使うツールの種類や入出力、エラーからの復帰手順を設計・見直します。文脈の情報量も整理し、作業完了率や再試行回数で改善を評価します。

  • エージェントのツールや入力形式の設計
  • ツールの結果と次の行動を明確にしたいとき
  • 安全な再試行と停止条件を定めたいとき
affaan-m/ECC27.7万2026年10月12日 更新

agent-introspection-debugging

無料日本語概要

AIエージェントの失敗や同じ操作の繰り返しを記録し、原因の切り分け、小さな復旧操作、結果の報告まで進める手順を示して、根拠のある再試行につなげるスキル。

  • 同じツール操作を繰り返す原因の調査
  • 会話の肥大化による品質低下の調査
  • ファイルパスや環境の食い違いの確認
affaan-m/ECC27.7万2026年10月12日 更新

agent-introspection-debugging

無料日本語概要

AIエージェントが失敗や同じ操作を繰り返す原因を、エラーと実行状況から整理します。小さな復旧操作を試し、結果と根拠を引き継げる報告にまとめるスキルです。

  • エージェントの連続失敗を診断したいとき
  • 同じツール操作のループ調査
  • 情報の蓄積による出力劣化の点検
affaan-m/ECC27.7万2026年10月5日 更新

affaan-m のスキルをすべて見る

このスキルの問題を報告する