本文へ移動
cccskills

「speech recognition」の検索結果

36 件 ・ 関連度順

概要と使いどころ

Azure Speech to Text REST API for short audio (Python). Use for simple speech recognition of audio files up to 60 seconds without the Speech SDK. Triggers: "speech to text REST", "short audio transcription", "speech recognition REST API", "STT REST", "recognize speech REST". DO NOT USE FOR: Long audio (>60 seconds), real-time streaming, batch transcription, custom speech models, speech translation. Use Speech SDK or Batch Transcription API instead.

日本語の概要は準備中です。原文の説明を表示しています。

microsoft/skills3,1012026年10月10日 更新

Transcribe speech to text using Apple's Speech framework. Use when implementing live microphone transcription with AVAudioEngine, recognizing recorded audio files, handling speech and microphone authorization, choosing on-device vs server-backed SFSpeechRecognizer behavior, or adopting SpeechAnalyzer, SpeechTranscriber, DictationTranscriber, AssetInventory, and async result streams on iOS 26+.

日本語の概要は準備中です。原文の説明を表示しています。

dpearson2699/swift-ios-skills1,1882026年8月1日 更新

apple-ml

無料日本語概要

Apple オンデバイス機械学習フレームワークリファレンス。 Core ML / Create ML / Vision / Natural Language / Speech。 MLModel, MLModelConfiguration, MLMultiArray, MLComputeUnits, MLImageClassifier, MLTextClassifier, MLDataTable, VNImageRequestHandler, VNRecognizeTextRequest, VNCoreMLRequest, NLTagger, NLLanguageRecognizer, NLEmbedding, SFSpeechRecognizer, SFSpeechAudioBufferRecognitionRequest, SFTranscription。

Fandhe-AI/agent-reference-skills42026年10月11日 更新

Transcribe speech to text using the Speech framework. Use when implementing live microphone transcription with AVAudioEngine, recognizing pre-recorded audio files, configuring on-device vs server-based recognition, handling authorization flows, or adopting the new SpeechAnalyzer API (iOS 26+) for modern async/await speech-to-text.

日本語の概要は準備中です。原文の説明を表示しています。

JordanCoin/ios-skills-collection62026年9月10日 更新

whisper

無料

OpenAI's general-purpose speech recognition model. Supports 99 languages, transcription, translation to English, and language identification. Six model sizes from tiny (39M params) to large (1550M params). Use for speech-to-text, podcast transcription, or multilingual audio processing. Best for robust, multilingual ASR.

日本語の概要は準備中です。原文の説明を表示しています。

davila7/claude-code-templates3.3万2026年10月11日 更新

whisper

無料

OpenAI's general-purpose speech recognition model. Supports 99 languages, transcription, translation to English, and language identification. Six model sizes from tiny (39M params) to large (1550M params). Use for speech-to-text, podcast transcription, or multilingual audio processing. Best for robust, multilingual ASR.

日本語の概要は準備中です。原文の説明を表示しています。

Orchestra-Research/AI-Research-SKILLs1.3万2026年10月11日 更新

whisper

無料

OpenAI's general-purpose speech recognition model. Supports 99 languages, transcription, translation to English, and language identification. Six model sizes from tiny (39M params) to large (1550M params). Use for speech-to-text, podcast transcription, or multilingual audio processing. Best for robust, multilingual ASR.

日本語の概要は準備中です。原文の説明を表示しています。

huang-sh/DeepScience42026年7月15日 更新

whisper

無料

OpenAI's general-purpose speech recognition model. Supports 99 languages, transcription, translation to English, and language identification. Six model sizes from tiny (39M params) to large (1550M params). Use for speech-to-text, podcast transcription, or multilingual audio processing. Best for robust, multilingual ASR.

日本語の概要は準備中です。原文の説明を表示しています。

Lord1Egypt/awesome-skill-forge22026年6月10日 更新

Azure Speech to Text REST API for short audio (Python). Use for simple speech recognition of audio files up to 60 seconds without the Speech SDK.

日本語の概要は準備中です。原文の説明を表示しています。

netbarros/psique62026年4月22日 更新

Azure Speech to Text REST API for short audio (Python). Use for simple speech recognition of audio files up to 60 seconds without the Speech SDK.

日本語の概要は準備中です。原文の説明を表示しています。

lucaspmarie-a11y/claude-skills-vault52026年6月7日 更新

Azure Speech to Text REST API for short audio (Python). Use for simple speech recognition of audio files up to 60 seconds without the Speech SDK.

日本語の概要は準備中です。原文の説明を表示しています。

phoroth/AGENTIC32026年8月8日 更新

Azure Speech to Text REST API for short audio (Python). Use for simple speech recognition of audio files up to 60 seconds without the Speech SDK.

日本語の概要は準備中です。原文の説明を表示しています。

nimoqup046-collab/agora32026年3月26日 更新

Azure Speech to Text REST API for short audio (Python). Use for simple speech recognition of audio files up to 60 seconds without the Speech SDK.

日本語の概要は準備中です。原文の説明を表示しています。

MMEHDI0606/ai-agent-foundation-template22026年5月21日 更新

Voice agents represent the frontier of AI interaction - humans speaking naturally with AI systems. The challenge isn't just speech recognition and synthesis, it's achieving natural conversation flow with sub-800ms latency while handling interruptions, background noise, and emotional nuance. This skill covers two architectures: speech-to-speech (OpenAI Realtime API, lowest latency, most natural) and pipeline (STT→LLM→TTS, more control, easier to debug). Key insight: latency is the constraint. Hu

日本語の概要は準備中です。原文の説明を表示しています。

davila7/claude-code-templates3.3万2026年10月11日 更新

Operate ASRT SpeechRecognition Chinese ASR data, acoustic models, pinyin language model, and serving clients.

日本語の概要は準備中です。原文の説明を表示しています。

VectorSpaceLab/AREX-Skill3322026年9月3日 更新

Produce, direct, integrate, and quality-control ElevenLabs text-to-speech for narration, character performance, multilingual media, long-form voiceover, and streamed or batch AI-content audio. Use when selecting ElevenLabs models or voices; preparing spoken text; controlling pronunciation, pacing, emotion, and code-switching; using TTS APIs; cloning or designing a voice with consent; repairing artifacts; mastering deliverables; or evaluating generated speech. Excludes music, sound-effect generation, conversational-agent design, and general speech recognition except transcription used to verify TTS.

日本語の概要は準備中です。原文の説明を表示しています。

calesthio/generative-media-skills1972026年7月14日 更新

whisper

無料

Modelo de reconhecimento de fala de propósito geral da OpenAI. Suporta 99 idiomas, transcrição, tradução para inglês e identificação de idioma. Seis tamanhos de modelo, de tiny (39M parâmetros) a large (1550M parâmetros). Use para speech-to-text, transcrição de podcasts ou processamento de áudio multilíngue. Melhor para ASR robusto e multilíngue.

日本語の概要は準備中です。原文の説明を表示しています。

artubss/SKILLS-CLAUDE-CODE112026年5月17日 更新

Load this skill whenever the project must support speech input or voice control — dictation, "click by voice" grammars, Dragon NaturallySpeaking, Voice Control (iOS/macOS), Voice Access (Android), or any interface where users activate controls by speaking. Under no circumstances let an aria-label override a visible label without keeping the visible text intact and in order. Absolutely always ensure the accessible name contains the visible label, and load alongside keyboard/SKILL.md since speech tools commonly emulate keyboard and pointer input.

日本語の概要は準備中です。原文の説明を表示しています。

LazyCats-dev/ao3-podfic-posting-helper112026年10月5日 更新

Voice agents represent the frontier of AI interaction - humans speaking naturally with AI systems. The challenge isn't just speech recognition and synthesis, it's achieving natural conversation flow with sub-800ms latency while handling interruptions, background noise, and emotional nuance. This skill covers two architectures: speech-to-speech (OpenAI Realtime API, lowest latency, most natural) and pipeline (STT→LLM→TTS, more control, easier to debug). Key insight: latency is the constraint. Hu

日本語の概要は準備中です。原文の説明を表示しています。

AxelMrak/ai52026年2月20日 更新

AssemblyAI is a hosted speech-to-text API that transcribes audio and video files or live streams and adds speaker labels, sentiment, entity detection, PII redaction and LLM analysis of the transcript. Use when a user asks to transcribe a recording, label who said what, analyze call sentiment, redact personal data, stream live transcription, or summarize and question a transcript with LLM Gateway (the replacement for LeMUR).

日本語の概要は準備中です。原文の説明を表示しています。

TerminalSkills/skills1632026年10月4日 更新

whisper

無料

OpenAI's general-purpose speech recognition model. Supports 99 languages, transcription, translation to English, and language identification. Six model sizes from tiny (39M params) to large (1550M params). Use for speech-to-text, podcast transcription, or multilingual audio processing. Best for robust, multilingual ASR.

日本語の概要は準備中です。原文の説明を表示しています。

ibragimov-oasis/vibe-coder22026年6月24日 更新

Use Transformers.js to run state-of-the-art machine learning models directly in JavaScript/TypeScript. Supports NLP (text classification, translation, summarization), computer vision (image classification, object detection), audio (speech recognition, audio classification), and multimodal tasks. Works in browsers and server-side runtimes (Node.js, Bun, Deno) with WebGPU/WASM using pre-trained models from Hugging Face Hub.

日本語の概要は準備中です。原文の説明を表示しています。

huggingface/skills1.1万2026年10月9日 更新

This skill should be used when working with pre-trained transformer models for natural language processing, computer vision, audio, or multimodal tasks. Use for text generation, classification, question answering, translation, summarization, image classification, object detection, speech recognition, and fine-tuning models on custom datasets.

日本語の概要は準備中です。原文の説明を表示しています。

foryourhealth111-pixel/Vibe-Skills3,6532026年8月31日 更新

Build call automation workflows with Azure Communication Services Call Automation Java SDK. Use when implementing IVR systems, call routing, call recording, DTMF recognition, text-to-speech, or AI-powered call flows.

日本語の概要は準備中です。原文の説明を表示しています。

microsoft/skills3,1012026年10月10日 更新