语音转文字 (ASR) (语音转文字 (ASR / 语音识别 / Speech-to-Text) 应用工程 (从业者视角) — 选型、集成、优化语音转文字能力,尤其面向移动端、低成本、快速识别的场景。覆盖: (a) 引擎/API 地图 — 云端 API(OpenAI gpt-4o-transcribe / Whisper API、Deepgram、AssemblyAI、Google / Azure / AWS Transcribe、讯飞、字节火山、阿里、腾讯、百度) vs 开源模型(Whisper / faster-whisper / whisper.cpp / distil-whisper、NVIDIA NeMo Parakeet / Canary、阿里 FunASR / Paraformer / SenseVoice、Moonshine、Vosk、Kaldi / k2 / icefall) vs 端侧·移动 SDK(whisper.cpp + CoreML / Metal、iOS Speech framework、Android SpeechRecognizer、Picovoice、SenseVoice 端侧); (b) 准确率(WER / CER) × 延迟(RTF / 流式) × 成本 三角权衡与选型决策树; (c) 成本优化 playbook — 端侧免费 / 批量折扣 / VAD 裁静音 / 量化(int8 / ggml) / 蒸馏 / 自托管 break-even; (d) 移动端集成 — 端侧 vs 云、流式 vs 批量、断点检测(endpointing)、隐私 / 离线; (e) 后处理 — 标点 / 数字规整(ITN) / 说话人分离(diarization) / 时间戳。学派分歧: 云 API vs 端侧自托管、通用大模型(Whisper) vs 专用流式(RNN-T / Conformer)、闭源 API vs 开源、准确率派 vs 成本派、英文优先 vs 中文 ASR(FunASR / SenseVoice / 讯飞)。不含: 文字转语音(TTS / 语音合成,方向相反)、声纹识别 / 说话人验证为主业、语音 agent / 对话式 AI、ASR 模型训练科研深水区。) Master OS — automated mastery of 语音转文字 (ASR / 语音识别 / Speech-to-Text) 应用工程 (从业者视角) — 选型、集成、优化语音转文字能力,尤其面向移动端、低成本、快速识别的场景。覆盖: (a) 引擎/API 地图 — 云端 API(OpenAI gpt-4o-transcribe / Whisper API、Deepgram、AssemblyAI、Google / Azure / AWS Transcribe、讯飞、字节火山、阿里、腾讯、百度) vs 开源模型(Whisper / faster-whisper / whisper.cpp / distil-whisper、NVIDIA NeMo Parakeet / Canary、阿里 FunASR / Paraformer / SenseVoice、Moonshine、Vosk、Kaldi / k2 / icefall) vs 端侧·移动 SDK(whisper.cpp + CoreML / Metal、iOS Speech framework、Android SpeechRecognizer、Picovoice、SenseVoice 端侧); (b) 准确率(WER / CER) × 延迟(RTF / 流式) × 成本 三角权衡与选型决策树; (c) 成本优化 playbook — 端侧免费 / 批量折扣 / VAD 裁静音 / 量化(int8 / ggml) / 蒸馏 / 自托管 break-even; (d) 移动端集成 — 端侧 vs 云、流式 vs 批量、断点检测(endpointing)、隐私 / 离线; (e) 后处理 — 标点 / 数字规整(ITN) / 说话人分离(diarization) / 时间戳。学派分歧: 云 API vs 端侧自托管、通用大模型(Whisper) vs 专用流式(RNN-T / Conformer)、闭源 API vs 开源、准确率派 vs 成本派、英文优先 vs 中文 ASR(FunASR / SenseVoice / 讯飞)。不含: 文字转语音(TTS / 语音合成,方向相反)、声纹识别 / 说话人验证为主业、语音 agent / 对话式 AI、ASR 模型训练科研深水区。: top builders' mental models, tool stack, current workflows, jargon, and where to keep up. Trigger this skill when the user works on 语音转文字 (ASR / 语音识别 / Speech-to-Text) 应用工程 (从业者视角) — 选型、集成、优化语音转文字能力,尤其面向移动端、低成本、快速识别的场景。覆盖: (a) 引擎/API 地图 — 云端 API(OpenAI gpt-4o-transcribe / Whisper API、Deepgram、AssemblyAI、Google / Azure / AWS Transcribe、讯飞、字节火山、阿里、腾讯、百度) vs 开源模型(Whisper / faster-whisper / whisper.cpp / distil-whisper、NVIDIA NeMo Parakeet / Canary、阿里 FunASR / Paraformer / SenseVoice、Moonshine、Vosk、Kaldi / k2 / icefall) vs 端侧·移动 SDK(whisper.cpp + CoreML / Metal、iOS Speech framework、Android SpeechRecognizer、Picovoice、SenseVoice 端侧); (b) 准确率(WER / CER) × 延迟(RTF / 流式) × 成本 三角权衡与选型决策树; (c) 成本优化 playbook — 端侧免费 / 批量折扣 / VAD 裁静音 / 量化(int8 / ggml) / 蒸馏 / 自托管 break-even; (d) 移动端集成 — 端侧 vs 云、流式 vs 批量、断点检测(endpointing)、隐私 / 离线; (e) 后处理 — 标点 / 数字规整(ITN) / 说话人分离(diarization) / 时间戳。学派分歧: 云 API vs 端侧自托管、通用大模型(Whisper) vs 专用流式(RNN-T / Conformer)、闭源 API vs 开源、准确率派 vs 成本派、英文优先 vs 中文 ASR(FunASR / SenseVoice / 讯飞)。不含: 文字转语音(TTS / 语音合成,方向相反)、声纹识别 / 说话人验证为主业、语音 agent / 对话式 AI、ASR 模型训练科研深水区。 problems and wants industry-grade thinking, tool selection, or workflow guidance. 触发词:「语音转文字」「语音识别」「asr」「speech to text」「stt」
日本語の概要は準備中です。原文の説明を表示しています。
swaylq/master-skill☆ 1482026年9月6日 更新
Quantizes LLMs to 8-bit or 4-bit for 50-75% memory reduction with minimal accuracy loss. Use when GPU memory is limited, need to fit larger models, or want faster inference. Supports INT8, NF4, FP4 formats, QLoRA training, and 8-bit optimizers. Works with HuggingFace Transformers.
日本語の概要は準備中です。原文の説明を表示しています。
davila7/claude-code-templates☆ 3.3万2026年10月11日 更新
Quantizes LLMs to 8-bit or 4-bit for 50-75% memory reduction with minimal accuracy loss. Use when GPU memory is limited, need to fit larger models, or want faster inference. Supports INT8, NF4, FP4 formats, QLoRA training, and 8-bit optimizers. Works with HuggingFace Transformers.
日本語の概要は準備中です。原文の説明を表示しています。
Orchestra-Research/AI-Research-SKILLs☆ 1.3万2026年6月16日 更新
把本地长视频/音频转写成文字稿 + 可选字幕,纯本地(不上传云端),用 sherpa-onnx X-ASR Zipformer transducer 模型(int8 量化、中英双语、自动标点)。已在 macOS Apple Silicon(int8 + AMX,~100× 实时)、Linux ARM64(CPU,~32× 实时)与 Windows(PowerShell 5.1,CPU)端到端验证。使用 `$lecture-to-md` 默认 ASR 后端;默认不要换成 Qwen。
日本語の概要は準備中です。原文の説明を表示しています。
ysyecust/lecture-to-notes☆ 2742026年10月3日 更新
Quantiza LLMs para 8-bit ou 4-bit com redução de memória de 50-75% e perda mínima de acurácia. Use quando a memória GPU é limitada, precisa ajustar modelos maiores ou quer inferência mais rápida. Suporta formatos INT8, NF4, FP4, treinamento QLoRA e otimizadores 8-bit. Funciona com HuggingFace Transformers.
日本語の概要は準備中です。原文の説明を表示しています。
artubss/SKILLS-CLAUDE-CODE☆ 112026年5月17日 更新
Quantizes LLMs to 8-bit or 4-bit for 50-75% memory reduction with minimal accuracy loss. Use when GPU memory is limited, need to fit larger models, or want faster inference. Supports INT8, NF4, FP4 formats, QLoRA training, and 8-bit optimizers. Works with HuggingFace Transformers.
日本語の概要は準備中です。原文の説明を表示しています。
huang-sh/DeepScience☆ 42026年7月15日 更新
Build MiniMax H3 (Hailuo) local video workflows with native T2V/I2V/R2V nodes, Comfy-Org INT8 weights, turbo LoRAs for 8GB VRAM, 15-second stereo-audio clips, and the official MiniMax prompting guides (cite by link, do not copy).
日本語の概要は準備中です。原文の説明を表示しています。
artokun/comfyui-mcp☆ 8072026年10月5日 更新
Routes DeepStream-Yolo deployment, model conversion, multi-GIE, and INT8 workflows for supported YOLO-family models.
日本語の概要は準備中です。原文の説明を表示しています。
VectorSpaceLab/AREX-Skill☆ 3312026年9月3日 更新
Think and work like an expert Edge / Embedded AI Engineer. Use when a task calls for Edge / Embedded AI Engineer judgment. Reasons from tensor-arena budgets, full-int8 PTQ with representative calibration, and TFLM/CMSIS-NN or Vela/Ethos-U compile paths through ONNX Runtime QNN HTP and mobile delegates—treating train–serve preprocessing skew, float thresholds on quantized outputs, and NPU operator fallback as first-class failure modes.
日本語の概要は準備中です。原文の説明を表示しています。
K-Dense-AI/scientific-agents☆ 2002026年10月3日 更新
Deploy machine learning models to edge devices using Google AI Edge Gallery, TensorFlow Lite, ONNX Runtime, and MediaPipe. Covers model quantization (INT8/INT4), on-device inference with Gemma 4 models, Android/iOS deployment via AI Edge Gallery, hardware delegate selection (GPU/NPU/DSP), and performance benchmarking on constrained devices. Use when deploying models to mobile phones, IoT devices, or embedded systems where cloud inference is impractical due to latency, cost, or connectivity constraints.
日本語の概要は準備中です。原文の説明を表示しています。
pjt222/agent-almanac☆ 372026年10月10日 更新
量化格式契约与位级对齐:从设备字节反推量化器的**编码公式、舍入模式、scale 粒度、退化块规则**, 并据此重实现/融合量化器(MXFP8 e8m0、int8 perblock、fp8 perchannel 等), 用「同进程同数据 + 独立参考实现 + 中点与退化输入全覆盖」做到**逐字节精确**。 当自研量化 kernel 与框架算子**对不上**、需要逆向某个量化器的数值契约、需要判断 「量化残差算不算可接受」、或要在融合内核里复现 aclnn 量化语义时使用此 skill。 即使用户只说「量化结果和参考不一致」「这个 scale 怎么算出来的」「融合量化器精度对不上」 而未提"契约",也应触发。**注意**:单纯选量化档位/配置(该不该开量化、开哪一档)走 dit-perf-opt;算子级性能调优与 DSL 选型走 operator-dev; 并行作用域导致的静默失效(DiT 侧「改了并行但不报错也没生效」的判定)走 dit-parallel-opt。
日本語の概要は準備中です。原文の説明を表示しています。
Ascend/MindIE-SD☆ 152026年10月10日 更新
在授权测试中核验 Node.js 权限模型(--permission / --experimental-permission,及旧 policy.json)能否被内部 API 绕过。当目标依赖 Node.js Permission Model 做沙箱隔离、需要判断给定 Node 版本下是否存在已知逃逸(inspector 断点、process.binding、fs.statfs/openAsBlob、process.mainModule.require、Module._load、path.resolve 覆盖、Uint8Array 路径、符号链接重命名)时使用。适用目标类型 Node.js 运行时 / 依赖权限模型的服务或 CLI。触发场景包括用户说"这个 Node 权限模型靠谱吗""--permission 能被绕过吗""测下 policy.json 沙箱逃逸""inspector/process.binding 能突破权限吗"。输出:给定 Node 版本下各绕过原语的命中矩阵 + 建议升级版本(含 killed 记录)。
日本語の概要は準備中です。原文の説明を表示しています。
galact-byte/galact-Skills☆ 72026年10月2日 更新