本文へ移動
cccskills

「tree attention」の検索結果

4 件 ・ 関連度順

概要と使いどころ

Accelerate LLM inference using speculative decoding, Medusa multiple heads, and lookahead decoding techniques. Use when optimizing inference speed (1.5-3.6× speedup), reducing latency for real-time applications, or deploying models with limited compute. Covers draft models, tree-based attention, Jacobi iteration, parallel token generation, and production deployment strategies.

日本語の概要は準備中です。原文の説明を表示しています。

davila7/claude-code-templates3.3万2026年10月11日 更新

Accelerate LLM inference using speculative decoding, Medusa multiple heads, and lookahead decoding techniques. Use when optimizing inference speed (1.5-3.6× speedup), reducing latency for real-time applications, or deploying models with limited compute. Covers draft models, tree-based attention, Jacobi iteration, parallel token generation, and production deployment strategies.

日本語の概要は準備中です。原文の説明を表示しています。

Orchestra-Research/AI-Research-SKILLs1.3万2026年6月16日 更新

Generate a prioritized review queue of registry skills and catalog items that need attention. Use this skill whenever someone asks: "what needs review?", "what's overdue for audit?", "where are the weak spots in the registry?", "what should I audit next?", "run a health check on the registry", "show me flagged skills", "find problems in the registry", "what skills need fixing?", or "triage the registry". This is the triage layer — it surfaces candidates ranked P0–P4 so that focused work (via /gaia-audit) targets the right things first. Run this before any audit pass. Also invoke proactively when a PR changes registry/named/ at scale, or before a release, to catch regressions early.

日本語の概要は準備中です。原文の説明を表示しています。

gaia-research/gaia-skill-tree232026年10月11日 更新

Acelere a inferência de LLMs usando especulative decoding, múltiplas cabeças Medusa e técnicas de lookahead decoding. Use ao otimizar velocidade de inferência (aceleração de 1,5-3,6×), reduzir latência em aplicações em tempo real ou fazer deploy de modelos com recursos computacionais limitados. Cobre modelos draft, atenção em árvore, iteração de Jacobi, geração paralela de tokens e estratégias de deploy em produção.

日本語の概要は準備中です。原文の説明を表示しています。

artubss/SKILLS-CLAUDE-CODE112026年5月17日 更新