Accelerate LLM inference using speculative decoding, Medusa multiple heads, and lookahead decoding techniques. Use when optimizing inference speed (1.5-3.6× speedup), reducing latency for real-time applications, or deploying models with limited compute. Covers draft models, tree-based attention, Jacobi iteration, parallel token generation, and production deployment strategies.
日本語の概要は準備中です。原文の説明を表示しています。
davila7/claude-code-templates☆ 3.3万2026年10月11日 更新
Accelerate LLM inference using speculative decoding, Medusa multiple heads, and lookahead decoding techniques. Use when optimizing inference speed (1.5-3.6× speedup), reducing latency for real-time applications, or deploying models with limited compute. Covers draft models, tree-based attention, Jacobi iteration, parallel token generation, and production deployment strategies.
日本語の概要は準備中です。原文の説明を表示しています。
Orchestra-Research/AI-Research-SKILLs☆ 1.3万2026年6月16日 更新
Generate a prioritized review queue of registry skills and catalog items that need attention. Use this skill whenever someone asks: "what needs review?", "what's overdue for audit?", "where are the weak spots in the registry?", "what should I audit next?", "run a health check on the registry", "show me flagged skills", "find problems in the registry", "what skills need fixing?", or "triage the registry". This is the triage layer — it surfaces candidates ranked P0–P4 so that focused work (via /gaia-audit) targets the right things first. Run this before any audit pass. Also invoke proactively when a PR changes registry/named/ at scale, or before a release, to catch regressions early.
日本語の概要は準備中です。原文の説明を表示しています。
gaia-research/gaia-skill-tree☆ 232026年10月11日 更新
Acelere a inferência de LLMs usando especulative decoding, múltiplas cabeças Medusa e técnicas de lookahead decoding. Use ao otimizar velocidade de inferência (aceleração de 1,5-3,6×), reduzir latência em aplicações em tempo real ou fazer deploy de modelos com recursos computacionais limitados. Cobre modelos draft, atenção em árvore, iteração de Jacobi, geração paralela de tokens e estratégias de deploy em produção.
日本語の概要は準備中です。原文の説明を表示しています。
artubss/SKILLS-CLAUDE-CODE☆ 112026年5月17日 更新