Azure Speech to Text REST API for short audio (Python). Use for simple speech recognition of audio files up to 60 seconds without the Speech SDK. Triggers: "speech to text REST", "short audio transcription", "speech recognition REST API", "STT REST", "recognize speech REST". DO NOT USE FOR: Long audio (>60 seconds), real-time streaming, batch transcription, custom speech models, speech translation. Use Speech SDK or Batch Transcription API instead.
日本語の概要は準備中です。原文の説明を表示しています。
microsoft/skills☆ 3,1012026年10月10日 更新
Transcribe speech to text using Apple's Speech framework. Use when implementing live microphone transcription with AVAudioEngine, recognizing recorded audio files, handling speech and microphone authorization, choosing on-device vs server-backed SFSpeechRecognizer behavior, or adopting SpeechAnalyzer, SpeechTranscriber, DictationTranscriber, AssetInventory, and async result streams on iOS 26+.
日本語の概要は準備中です。原文の説明を表示しています。
dpearson2699/swift-ios-skills☆ 1,1882026年8月1日 更新
Apple オンデバイス機械学習フレームワークリファレンス。 Core ML / Create ML / Vision / Natural Language / Speech。 MLModel, MLModelConfiguration, MLMultiArray, MLComputeUnits, MLImageClassifier, MLTextClassifier, MLDataTable, VNImageRequestHandler, VNRecognizeTextRequest, VNCoreMLRequest, NLTagger, NLLanguageRecognizer, NLEmbedding, SFSpeechRecognizer, SFSpeechAudioBufferRecognitionRequest, SFTranscription。
Fandhe-AI/agent-reference-skills☆ 42026年10月11日 更新
Transcribe speech to text using the Speech framework. Use when implementing live microphone transcription with AVAudioEngine, recognizing pre-recorded audio files, configuring on-device vs server-based recognition, handling authorization flows, or adopting the new SpeechAnalyzer API (iOS 26+) for modern async/await speech-to-text.
日本語の概要は準備中です。原文の説明を表示しています。
JordanCoin/ios-skills-collection☆ 62026年9月10日 更新
OpenAI's general-purpose speech recognition model. Supports 99 languages, transcription, translation to English, and language identification. Six model sizes from tiny (39M params) to large (1550M params). Use for speech-to-text, podcast transcription, or multilingual audio processing. Best for robust, multilingual ASR.
日本語の概要は準備中です。原文の説明を表示しています。
davila7/claude-code-templates☆ 3.3万2026年10月11日 更新
OpenAI's general-purpose speech recognition model. Supports 99 languages, transcription, translation to English, and language identification. Six model sizes from tiny (39M params) to large (1550M params). Use for speech-to-text, podcast transcription, or multilingual audio processing. Best for robust, multilingual ASR.
日本語の概要は準備中です。原文の説明を表示しています。
Orchestra-Research/AI-Research-SKILLs☆ 1.3万2026年10月11日 更新
OpenAI's general-purpose speech recognition model. Supports 99 languages, transcription, translation to English, and language identification. Six model sizes from tiny (39M params) to large (1550M params). Use for speech-to-text, podcast transcription, or multilingual audio processing. Best for robust, multilingual ASR.
日本語の概要は準備中です。原文の説明を表示しています。
huang-sh/DeepScience☆ 42026年7月15日 更新
OpenAI's general-purpose speech recognition model. Supports 99 languages, transcription, translation to English, and language identification. Six model sizes from tiny (39M params) to large (1550M params). Use for speech-to-text, podcast transcription, or multilingual audio processing. Best for robust, multilingual ASR.
日本語の概要は準備中です。原文の説明を表示しています。
Lord1Egypt/awesome-skill-forge☆ 22026年6月10日 更新
Azure Speech to Text REST API for short audio (Python). Use for simple speech recognition of audio files up to 60 seconds without the Speech SDK.
日本語の概要は準備中です。原文の説明を表示しています。
netbarros/psique☆ 62026年4月22日 更新
Azure Speech to Text REST API for short audio (Python). Use for simple speech recognition of audio files up to 60 seconds without the Speech SDK.
日本語の概要は準備中です。原文の説明を表示しています。
lucaspmarie-a11y/claude-skills-vault☆ 52026年6月7日 更新
Azure Speech to Text REST API for short audio (Python). Use for simple speech recognition of audio files up to 60 seconds without the Speech SDK.
日本語の概要は準備中です。原文の説明を表示しています。
phoroth/AGENTIC☆ 32026年8月8日 更新
Azure Speech to Text REST API for short audio (Python). Use for simple speech recognition of audio files up to 60 seconds without the Speech SDK.
日本語の概要は準備中です。原文の説明を表示しています。
nimoqup046-collab/agora☆ 32026年3月26日 更新
Azure Speech to Text REST API for short audio (Python). Use for simple speech recognition of audio files up to 60 seconds without the Speech SDK.
日本語の概要は準備中です。原文の説明を表示しています。
MMEHDI0606/ai-agent-foundation-template☆ 22026年5月21日 更新
Voice agents represent the frontier of AI interaction - humans speaking naturally with AI systems. The challenge isn't just speech recognition and synthesis, it's achieving natural conversation flow with sub-800ms latency while handling interruptions, background noise, and emotional nuance. This skill covers two architectures: speech-to-speech (OpenAI Realtime API, lowest latency, most natural) and pipeline (STT→LLM→TTS, more control, easier to debug). Key insight: latency is the constraint. Hu
日本語の概要は準備中です。原文の説明を表示しています。
davila7/claude-code-templates☆ 3.3万2026年10月11日 更新
Operate ASRT SpeechRecognition Chinese ASR data, acoustic models, pinyin language model, and serving clients.
日本語の概要は準備中です。原文の説明を表示しています。
VectorSpaceLab/AREX-Skill☆ 3322026年9月3日 更新
Produce, direct, integrate, and quality-control ElevenLabs text-to-speech for narration, character performance, multilingual media, long-form voiceover, and streamed or batch AI-content audio. Use when selecting ElevenLabs models or voices; preparing spoken text; controlling pronunciation, pacing, emotion, and code-switching; using TTS APIs; cloning or designing a voice with consent; repairing artifacts; mastering deliverables; or evaluating generated speech. Excludes music, sound-effect generation, conversational-agent design, and general speech recognition except transcription used to verify TTS.
日本語の概要は準備中です。原文の説明を表示しています。
calesthio/generative-media-skills☆ 1972026年7月14日 更新
Modelo de reconhecimento de fala de propósito geral da OpenAI. Suporta 99 idiomas, transcrição, tradução para inglês e identificação de idioma. Seis tamanhos de modelo, de tiny (39M parâmetros) a large (1550M parâmetros). Use para speech-to-text, transcrição de podcasts ou processamento de áudio multilíngue. Melhor para ASR robusto e multilíngue.
日本語の概要は準備中です。原文の説明を表示しています。
artubss/SKILLS-CLAUDE-CODE☆ 112026年5月17日 更新
Load this skill whenever the project must support speech input or voice control — dictation, "click by voice" grammars, Dragon NaturallySpeaking, Voice Control (iOS/macOS), Voice Access (Android), or any interface where users activate controls by speaking. Under no circumstances let an aria-label override a visible label without keeping the visible text intact and in order. Absolutely always ensure the accessible name contains the visible label, and load alongside keyboard/SKILL.md since speech tools commonly emulate keyboard and pointer input.
日本語の概要は準備中です。原文の説明を表示しています。
LazyCats-dev/ao3-podfic-posting-helper☆ 112026年10月5日 更新
Voice agents represent the frontier of AI interaction - humans speaking naturally with AI systems. The challenge isn't just speech recognition and synthesis, it's achieving natural conversation flow with sub-800ms latency while handling interruptions, background noise, and emotional nuance. This skill covers two architectures: speech-to-speech (OpenAI Realtime API, lowest latency, most natural) and pipeline (STT→LLM→TTS, more control, easier to debug). Key insight: latency is the constraint. Hu
日本語の概要は準備中です。原文の説明を表示しています。
AxelMrak/ai☆ 52026年2月20日 更新
AssemblyAI is a hosted speech-to-text API that transcribes audio and video files or live streams and adds speaker labels, sentiment, entity detection, PII redaction and LLM analysis of the transcript. Use when a user asks to transcribe a recording, label who said what, analyze call sentiment, redact personal data, stream live transcription, or summarize and question a transcript with LLM Gateway (the replacement for LeMUR).
日本語の概要は準備中です。原文の説明を表示しています。
TerminalSkills/skills☆ 1632026年10月4日 更新
OpenAI's general-purpose speech recognition model. Supports 99 languages, transcription, translation to English, and language identification. Six model sizes from tiny (39M params) to large (1550M params). Use for speech-to-text, podcast transcription, or multilingual audio processing. Best for robust, multilingual ASR.
日本語の概要は準備中です。原文の説明を表示しています。
ibragimov-oasis/vibe-coder☆ 22026年6月24日 更新
Use Transformers.js to run state-of-the-art machine learning models directly in JavaScript/TypeScript. Supports NLP (text classification, translation, summarization), computer vision (image classification, object detection), audio (speech recognition, audio classification), and multimodal tasks. Works in browsers and server-side runtimes (Node.js, Bun, Deno) with WebGPU/WASM using pre-trained models from Hugging Face Hub.
日本語の概要は準備中です。原文の説明を表示しています。
huggingface/skills☆ 1.1万2026年10月9日 更新
This skill should be used when working with pre-trained transformer models for natural language processing, computer vision, audio, or multimodal tasks. Use for text generation, classification, question answering, translation, summarization, image classification, object detection, speech recognition, and fine-tuning models on custom datasets.
日本語の概要は準備中です。原文の説明を表示しています。
foryourhealth111-pixel/Vibe-Skills☆ 3,6532026年8月31日 更新
Build call automation workflows with Azure Communication Services Call Automation Java SDK. Use when implementing IVR systems, call routing, call recording, DTMF recognition, text-to-speech, or AI-powered call flows.
日本語の概要は準備中です。原文の説明を表示しています。
microsoft/skills☆ 3,1012026年10月10日 更新