本文へ移動
cccskills

「latency reduction」の検索結果

8 件 ・ 関連度順

概要と使いどころ

Reduces LLM costs and improves response times through caching, model selection, batching, and prompt optimization. Provides cost breakdowns, latency hotspots, and configuration recommendations. Use for "cost reduction", "performance optimization", "latency improvement", or "efficiency".

日本語の概要は準備中です。原文の説明を表示しています。

sathishssj3/Stereix-Engine22026年10月4日 更新

Accelerate LLM inference using speculative decoding, Medusa multiple heads, and lookahead decoding techniques. Use when optimizing inference speed (1.5-3.6× speedup), reducing latency for real-time applications, or deploying models with limited compute. Covers draft models, tree-based attention, Jacobi iteration, parallel token generation, and production deployment strategies.

日本語の概要は準備中です。原文の説明を表示しています。

davila7/claude-code-templates3.3万2026年10月11日 更新

Provides recommendations for environmental sustainability, carbon footprint reduction, and energy efficiency based on the Sustainability pillar of the Google Cloud Well-Architected Framework (WAF). Use when the user asks to assess, design, or optimize Google Cloud workloads for sustainability—including the shared responsibility model, selecting low-carbon regions (CFE%), reducing resource and AI/ML energy waste, designing efficient software and storage lifecycles, or measuring and tracking emissions using Google Cloud Carbon Footprint. Don't use for financial cost reduction (use google-cloud-waf-cost-optimization), latency and throughput tuning (use google-cloud-waf-performance-optimization), or high availability and disaster recovery (use google-cloud-waf-reliability).

日本語の概要は準備中です。原文の説明を表示しています。

google/skills2.1万2026年10月10日 更新

Accelerate LLM inference using speculative decoding, Medusa multiple heads, and lookahead decoding techniques. Use when optimizing inference speed (1.5-3.6× speedup), reducing latency for real-time applications, or deploying models with limited compute. Covers draft models, tree-based attention, Jacobi iteration, parallel token generation, and production deployment strategies.

日本語の概要は準備中です。原文の説明を表示しています。

Orchestra-Research/AI-Research-SKILLs1.3万2026年10月11日 更新

Optimize Claude API performance with prompt caching, model selection, streaming, and latency reduction techniques. Use when experiencing slow responses, optimizing token usage, or reducing time-to-first-token in production. Trigger with phrases like "anthropic performance", "claude speed", "optimize claude latency", "anthropic caching", "faster claude responses".

日本語の概要は準備中です。原文の説明を表示しています。

jeremylongshore/tons-of-skills-marketplace2,8312026年10月11日 更新

Use when you need to refactor Java code for high performance — including memory/allocation reduction, CPU hot-path optimization, and syntax/API/control-flow improvements. This should trigger for requests such as Review Java code for high performance; Optimize Java hot path; Reduce Java allocations; Improve Java latency/throughput. Part of Plinth Toolkit

日本語の概要は準備中です。原文の説明を表示しています。

jabrena/plinth4482026年10月8日 更新

Acelere a inferência de LLMs usando especulative decoding, múltiplas cabeças Medusa e técnicas de lookahead decoding. Use ao otimizar velocidade de inferência (aceleração de 1,5-3,6×), reduzir latência em aplicações em tempo real ou fazer deploy de modelos com recursos computacionais limitados. Cobre modelos draft, atenção em árvore, iteração de Jacobi, geração paralela de tokens e estratégias de deploy em produção.

日本語の概要は準備中です。原文の説明を表示しています。

artubss/SKILLS-CLAUDE-CODE112026年5月17日 更新

tw-reduce

無料

Audit over-engineered codebases by factoring layers into live obligations, quotienting redundant distinctions, ablating unearned surface, and normalizing survivors while preserving required behavior. Use when change latency or agent difficulty comes from frameworks, plugins, DI, codegen, task runners, config indirection, ORMs, GraphQL, monorepo/infra tooling, or web stacks, or when asked to remove layers. Produces an evidence-backed Reduction Certificate, cuts, migration phases, proof signals, and rollback. Not for one local readability cleanup (use tw-complexity-mitigator).

日本語の概要は準備中です。原文の説明を表示しています。

SDiamante13/dotfiles82026年10月2日 更新