本文へ移動
cccskills

「whisper」の検索結果

110 件 ・ 関連度順

概要と使いどころ

语音转文字 (ASR) (语音转文字 (ASR / 语音识别 / Speech-to-Text) 应用工程 (从业者视角) — 选型、集成、优化语音转文字能力,尤其面向移动端、低成本、快速识别的场景。覆盖: (a) 引擎/API 地图 — 云端 API(OpenAI gpt-4o-transcribe / Whisper API、Deepgram、AssemblyAI、Google / Azure / AWS Transcribe、讯飞、字节火山、阿里、腾讯、百度) vs 开源模型(Whisper / faster-whisper / whisper.cpp / distil-whisper、NVIDIA NeMo Parakeet / Canary、阿里 FunASR / Paraformer / SenseVoice、Moonshine、Vosk、Kaldi / k2 / icefall) vs 端侧·移动 SDK(whisper.cpp + CoreML / Metal、iOS Speech framework、Android SpeechRecognizer、Picovoice、SenseVoice 端侧); (b) 准确率(WER / CER) × 延迟(RTF / 流式) × 成本 三角权衡与选型决策树; (c) 成本优化 playbook — 端侧免费 / 批量折扣 / VAD 裁静音 / 量化(int8 / ggml) / 蒸馏 / 自托管 break-even; (d) 移动端集成 — 端侧 vs 云、流式 vs 批量、断点检测(endpointing)、隐私 / 离线; (e) 后处理 — 标点 / 数字规整(ITN) / 说话人分离(diarization) / 时间戳。学派分歧: 云 API vs 端侧自托管、通用大模型(Whisper) vs 专用流式(RNN-T / Conformer)、闭源 API vs 开源、准确率派 vs 成本派、英文优先 vs 中文 ASR(FunASR / SenseVoice / 讯飞)。不含: 文字转语音(TTS / 语音合成,方向相反)、声纹识别 / 说话人验证为主业、语音 agent / 对话式 AI、ASR 模型训练科研深水区。) Master OS — automated mastery of 语音转文字 (ASR / 语音识别 / Speech-to-Text) 应用工程 (从业者视角) — 选型、集成、优化语音转文字能力,尤其面向移动端、低成本、快速识别的场景。覆盖: (a) 引擎/API 地图 — 云端 API(OpenAI gpt-4o-transcribe / Whisper API、Deepgram、AssemblyAI、Google / Azure / AWS Transcribe、讯飞、字节火山、阿里、腾讯、百度) vs 开源模型(Whisper / faster-whisper / whisper.cpp / distil-whisper、NVIDIA NeMo Parakeet / Canary、阿里 FunASR / Paraformer / SenseVoice、Moonshine、Vosk、Kaldi / k2 / icefall) vs 端侧·移动 SDK(whisper.cpp + CoreML / Metal、iOS Speech framework、Android SpeechRecognizer、Picovoice、SenseVoice 端侧); (b) 准确率(WER / CER) × 延迟(RTF / 流式) × 成本 三角权衡与选型决策树; (c) 成本优化 playbook — 端侧免费 / 批量折扣 / VAD 裁静音 / 量化(int8 / ggml) / 蒸馏 / 自托管 break-even; (d) 移动端集成 — 端侧 vs 云、流式 vs 批量、断点检测(endpointing)、隐私 / 离线; (e) 后处理 — 标点 / 数字规整(ITN) / 说话人分离(diarization) / 时间戳。学派分歧: 云 API vs 端侧自托管、通用大模型(Whisper) vs 专用流式(RNN-T / Conformer)、闭源 API vs 开源、准确率派 vs 成本派、英文优先 vs 中文 ASR(FunASR / SenseVoice / 讯飞)。不含: 文字转语音(TTS / 语音合成,方向相反)、声纹识别 / 说话人验证为主业、语音 agent / 对话式 AI、ASR 模型训练科研深水区。: top builders' mental models, tool stack, current workflows, jargon, and where to keep up. Trigger this skill when the user works on 语音转文字 (ASR / 语音识别 / Speech-to-Text) 应用工程 (从业者视角) — 选型、集成、优化语音转文字能力,尤其面向移动端、低成本、快速识别的场景。覆盖: (a) 引擎/API 地图 — 云端 API(OpenAI gpt-4o-transcribe / Whisper API、Deepgram、AssemblyAI、Google / Azure / AWS Transcribe、讯飞、字节火山、阿里、腾讯、百度) vs 开源模型(Whisper / faster-whisper / whisper.cpp / distil-whisper、NVIDIA NeMo Parakeet / Canary、阿里 FunASR / Paraformer / SenseVoice、Moonshine、Vosk、Kaldi / k2 / icefall) vs 端侧·移动 SDK(whisper.cpp + CoreML / Metal、iOS Speech framework、Android SpeechRecognizer、Picovoice、SenseVoice 端侧); (b) 准确率(WER / CER) × 延迟(RTF / 流式) × 成本 三角权衡与选型决策树; (c) 成本优化 playbook — 端侧免费 / 批量折扣 / VAD 裁静音 / 量化(int8 / ggml) / 蒸馏 / 自托管 break-even; (d) 移动端集成 — 端侧 vs 云、流式 vs 批量、断点检测(endpointing)、隐私 / 离线; (e) 后处理 — 标点 / 数字规整(ITN) / 说话人分离(diarization) / 时间戳。学派分歧: 云 API vs 端侧自托管、通用大模型(Whisper) vs 专用流式(RNN-T / Conformer)、闭源 API vs 开源、准确率派 vs 成本派、英文优先 vs 中文 ASR(FunASR / SenseVoice / 讯飞)。不含: 文字转语音(TTS / 语音合成,方向相反)、声纹识别 / 说话人验证为主业、语音 agent / 对话式 AI、ASR 模型训练科研深水区。 problems and wants industry-grade thinking, tool selection, or workflow guidance. 触发词:「语音转文字」「语音识别」「asr」「speech to text」「stt」

日本語の概要は準備中です。原文の説明を表示しています。

swaylq/master-skill1492026年9月6日 更新

openai-whisper

無料日本語概要

音声ファイルをローカルで文字起こしし、テキストや翻訳字幕に出力するスキル。APIキー不要で、モデルの大きさを選び、処理速度と精度を調整できます。

  • 録音音声を文字起こししたいとき
  • 音声から翻訳字幕を作りたいとき
  • 速度や精度に合わせたモデル選び
openclaw/openclaw39.2万2026年10月11日 更新

local-media-transcription

無料日本語概要

Transcribe local MP4/M4A/MP3/WAV/WEBM audio or video with ffmpeg and Whisper, then optionally produce diarized transcripts, customer-facing meeting minutes, action items, and PPT-ready summaries. Use when: 文字起こし, transcription, transcribe mp4, meeting transcript, mp4から文字起こし, 録画から議事録, 顧客向け議事録, 話者分離, diarization, whisper CLI, ffmpeg.

aktsmm/Agent-Skills262026年10月10日 更新

premiere-skills

無料日本語概要

Premiere Pro 動画編集ワークフローを Claude Code で自動化するスキル集。/cut は Premiere Pro XML の無音・雑音区間をジェットカット、/srt は WAV+Premiere XML から日本語テロップ用 SRT を並列Whisper→LLM意味区切り→全体アライメントSRT の3ステップ(v6)で生成、/srt-fast は改行工程まで3チャンク並列化した高速版、/telop-check は書き出し済みMP4を全編スキャンしてテロップの数字表記・誤字脱字・い抜き・ら抜き・固有名詞ミスを検出する。faster-whisper / ffmpeg / Premiere Pro を要するローカル実行型。日本語トーク動画のショート / ロング編集向け。

fuuuuuuma/premiere-skills142026年9月25日 更新

Local speech-to-text using faster-whisper. 4-6x faster than OpenAI Whisper with identical accuracy; GPU acceleration enables ~20x realtime transcription. Supports standard and distilled models with word-level timestamps.

日本語の概要は準備中です。原文の説明を表示しています。

danstrem2/clawdbot-skill-master-pack22026年2月1日 更新

openai-whisper-api

無料日本語概要

音声ファイルをOpenAIの文字起こしAPIに送り、テキストやJSONで保存します。モデルや言語の指定、話者を区別した文字起こしにも対応するスキル。

  • 会議の録音を文字起こししたいとき
  • インタビューの話者を区別したいとき
  • 文字起こしをJSONで保存したいとき
openclaw/openclaw39.2万2026年10月11日 更新

When you want to extract content from a video — YouTube, Loom, Vimeo, Riverside, Zoom recording, local MP4, X/IG video, anything yt-dlp supports. Three depth modes user picks per invocation — transcript (just words, fast/free), visual (transcript + ffmpeg frame extraction + Claude vision pass on key moments), multimodal (Gemini native video ingestion if $GEMINI_API_KEY set, else dense Claude vision). Uses local Whisper for transcription (MLX-Whisper on Apple Silicon, faster-whisper elsewhere), falls back to platform-provided transcripts when available (Loom, Riverside, YouTube auto-subs). Saves to ~/Documents/videos/<source>-<slug>-<date>/ and optionally captures summary to second-brain raw/ as call-/meeting-/note-. Triggers on "/watch-video <url>," "watch this video," "transcribe this loom," "analyze this video," "summarize this recording," "key moments from this," "what happened in this video." This skill replaces and broadens the prior youtube-transcript skill.

日本語の概要は準備中です。原文の説明を表示しています。

coreyhaines31/makerskills8542026年10月9日 更新

Use the faster-whisper package for CTranslate2-backed Whisper transcription, model selection, CPU/CUDA setup, audio utilities, VAD, timestamps, and conversion guidance.

日本語の概要は準備中です。原文の説明を表示しています。

VectorSpaceLab/AREX-Skill3312026年9月3日 更新

Use when the user wants local voice transcription instead of OpenAI Whisper API. Switches to whisper.cpp running on Apple Silicon. WhatsApp only for now. Requires voice-transcription skill to be applied first.

日本語の概要は準備中です。原文の説明を表示しています。

nanocoai/nanoclaw-skills182026年3月30日 更新

Use when the user wants local voice transcription instead of OpenAI Whisper API. Switches to whisper.cpp running on Apple Silicon. WhatsApp only for now. Requires voice-transcription skill to be applied first.

日本語の概要は準備中です。原文の説明を表示しています。

nanocoai/nanoclaw-whatsapp102026年4月15日 更新

Use when the user wants local voice transcription instead of OpenAI Whisper API. Switches to whisper.cpp running on Apple Silicon. WhatsApp only for now. Requires voice-transcription skill to be applied first.

日本語の概要は準備中です。原文の説明を表示しています。

nanocoai/nanoclaw-telegram102026年4月3日 更新

Transcribes audio/video files to text. Uses Whisper (openai-whisper) or Vosk (offline) as optional backend — both are detected via presence check. Without backend: placeholder mode with dummy output (dry-run).

日本語の概要は準備中です。原文の説明を表示しています。

ellmos-ai/skills72026年10月10日 更新

Use when the user wants local voice transcription instead of OpenAI Whisper API. Switches to whisper.cpp running on Apple Silicon. WhatsApp only for now. Requires voice-transcription skill to be applied first.

日本語の概要は準備中です。原文の説明を表示しています。

nanocoai/nanoclaw-gmail22026年4月3日 更新

whisper

無料

OpenAI's general-purpose speech recognition model. Supports 99 languages, transcription, translation to English, and language identification. Six model sizes from tiny (39M params) to large (1550M params). Use for speech-to-text, podcast transcription, or multilingual audio processing. Best for robust, multilingual ASR.

日本語の概要は準備中です。原文の説明を表示しています。

davila7/claude-code-templates3.3万2026年10月11日 更新

watch

無料

Watch a video (URL or local path). Downloads with yt-dlp, extracts auto-scaled frames with ffmpeg, pulls the transcript from captions (or local WhisperX / cloud Whisper fallback), and hands the result to the agent so it can answer questions about what's in the video. With a Gemini API key, Google's agentic video model watches the full video instead.

日本語の概要は準備中です。原文の説明を表示しています。

bradautomates/claude-video1.8万2026年9月25日 更新

whisper

無料

OpenAI's general-purpose speech recognition model. Supports 99 languages, transcription, translation to English, and language identification. Six model sizes from tiny (39M params) to large (1550M params). Use for speech-to-text, podcast transcription, or multilingual audio processing. Best for robust, multilingual ASR.

日本語の概要は準備中です。原文の説明を表示しています。

Orchestra-Research/AI-Research-SKILLs1.3万2026年6月16日 更新

subtitles

無料

Burn timed subtitles onto a finished video or configure Whisper-timed caption burning during faceless-video assembly. Takes video/audio generation jobs or a local finished video, optionally with authored narration text, and returns a captioned video or the exact backend subtitle configuration. Timings always come from Whisper on the video's own audio — never estimated. THREE looks: `paper` (torn cream paper label, handwritten ink), `bold` (UGC ALL-CAPS, white with black stroke, platform safe zones), `clean` (slim white CAPS, no plate, tiny at the bottom — no Pillow needed). Use when: adding/burning subtitles or captions to an existing video, restyling captions, or when a production workflow needs its finished cut subtitled. NOT for: generating the video, translating speech, or live/soft subtitle tracks (this burns them into the picture).

日本語の概要は準備中です。原文の説明を表示しています。

openai/plugins7,3872026年10月8日 更新

Local speech-to-text with the Whisper CLI (no API key).

日本語の概要は準備中です。原文の説明を表示しています。

huangruiteng/CS-Notes4,0022026年10月9日 更新

Transcribe audio segments to text using Whisper models. Use larger models (small, base, medium, large-v3) for better accuracy, or faster-whisper for optimized performance. Always align transcription timestamps with diarization segments for accurate speaker-labeled subtitles.

日本語の概要は準備中です。原文の説明を表示しています。

benchflow-ai/skillsbench1,8372026年7月24日 更新

Transcribe audio/video to text with word-level timestamps using OpenAI Whisper. Use when you need speech-to-text with accurate timing information for each word.

日本語の概要は準備中です。原文の説明を表示しています。

benchflow-ai/skillsbench1,8372026年7月24日 更新

Assemble a myth-vs-fact kinetic-typography explainer video ad (≈29.5s, 9:16) from N myth/fact pairs + hook / turn / punch copy + palette + a brand end-card PNG + a VO track — a hook, 3 red-strike MYTH cards that flip to teal-check FACT cards (per-line strikethrough that crosses EVERY wrapped line), a "what actually works" turn, an optional proof reveal, a punch line, and a static end card. DETERMINISTIC assembly with ZERO AI-gen visuals — HTML hyperframes rendered frame-exact via Playwright (`window.renderAt(t)`, animation a pure function of beat-local time), Whisper beat-snap to VO word onsets, concat at a uniform fps, karaoke `.ass` captions burned last (suppressed on the proof + end-card beats), and a VO + optional music mix (music −20 dB, `amix normalize=0`, tail fade). FREE (Python + Playwright + ffmpeg); the recipe supplies the copy / palette / end-card / VO and gates the paid VO / music / Whisper calls to their own capabilities. Use for the myth-vs-fact format.

日本語の概要は準備中です。原文の説明を表示しています。

gooseworks-ai/goose-skills1,2422026年10月10日 更新

Assemble a narrated-UGC "stitch reply" ad from a config — a single spoken VO carries a verbatim testimonial while ~30 per-cut i2v clips (one creator across ~5 wardrobes in ~3 worlds, plus product B-roll) are each trimmed to their EDL window built from the VO's Whisper word boundaries and hard-concatenated via filter_complex concat (never the demuxer, which drops audio on a duration mismatch), the VO mixed over an optional sidechain-ducked instrumental bed (−20dB, 20 to 1) so the VO stays on top, karaoke-pop captions burned on every word throughout (VEED Whisper preset, re-spelled against the locked script), a landing-page scroll rendered as FFmpeg zoompan over a Playwright PNG (not i2v), and closed on the brand's real end-card PNG — never AI-rendered text. This is the FREE deterministic assembly stage (trim-to-EDL + filter_complex concat + VO and music mix + karaoke captions + landing-page zoompan + end-card append), run with the shared stitch-videos-ffmpeg montage.py helper that installs with this package; the VO, creator, start-frames, and clips come from create-vo-elevenlabs / create-image-gpt-image-fal / create-image-fal / create-video-fal. Use for the narrated-ugc-wardrobe-stitch format.

日本語の概要は準備中です。原文の説明を表示しています。

gooseworks-ai/goose-skills1,2422026年10月10日 更新

Speech-to-text transcription via OpenAI Whisper. Supports two modes — Local CLI (no API key, runs on-device) and Cloud API (fast, scalable, requires OPENAI_API_KEY). Use when the user needs to transcribe audio files, translate speech, or convert audio to text.

日本語の概要は準備中です。原文の説明を表示しています。

coco-research/coco5362026年10月11日 更新

Routes Distil-Whisper inference, PyTorch distillation training, and Flax reproduction workflows.

日本語の概要は準備中です。原文の説明を表示しています。

VectorSpaceLab/AREX-Skill3312026年9月3日 更新