本文へ移動
cccskills
無料GitHub で公開

context-optimization

Use when optimizing token usage, KV cache efficiency, or context window management for LLM agents. Keywords: context optimization, KV cache, prompt caching, token budget, semantic pruning, lost-in-the-middle.

インストール方法を見る

含まれるファイル(1)

  • SKILL.md3.9 KB

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

Context Optimization Skill

This skill provides methodologies and best practices for maximizing context window efficiency, reducing token costs, and improving the performance of LLM-based agentic workflows.

1. Prompt Cache Alignment

To maximize Key-Value (KV) cache hits, prioritize stability in the early parts of the prompt.

  • Static Prefixes: Design system prompts to remain identical across agent turns.
  • Tool Order: Sort tool definitions alphabetically or by frequency of use. Keep this ordering constant.
  • Prefix Consistency: Reserve the top 70% of the KV cache for global system instructions and tool definitions.
  • Avoid Dynamic Data: Move dynamic session-specific data (dates, current file list) to the end of the context window.

2. Semantic Pruning & AST Compaction

Reduce unnecessary data before sending it to the model.

  • Removal of Non-Essentials: Strip comments, debug logs, and unused boilerplate code.
  • AST Compaction: Convert deep code structures into simplified, representative summaries or pseudocode.
  • Whitespace Stripping: Use minification for configuration files (JSON, YAML) and standard code.
  • Irrelevant Body Dropping: Replace large, irrelevant function bodies with docstrings or signatures if the model doesn't need to reason about the implementation details.

3. Lost-in-the-Middle Mitigation

LLMs often suffer from recall degradation for information in the center of the context window.

  • Boundary Priority: Place high-priority constraints, critical instructions, and schema definitions at the very beginning (Head) or the very end (Tail) of the context.
  • Sandwich Strategy: If critical information must be in the middle, sandwich it between two clear, high-level summary points that repeat its purpose.

4. Sliding Windows & Recursive Summarization

Manage long-running conversations without exceeding token thresholds.

  • Sliding Window: Keep only the N most recent turns for immediate interaction.
  • Recursive Summarization: Periodically collapse historical turns into a compressed "Session State Summary".
  • State Archiving: Store older, less relevant interactions in a side-car file (e.g., archive.md) that the agent can read only when necessary.

5. Token Budget Allocation

Implement a disciplined token distribution:

CategoryTypical %Goal
System Prompt10-15%Definition of role & constraints
Tool Definitions15-20%Capability exposure
Interaction History30-40%Context for current turn
Reserved (Drafting)25-45%Headroom for generation

6. Context Freshness Decisions

Determine when to act on history:

  • Summarize: When history exceeds 50% of the token budget and requires long-term context.
  • Truncate: When history is irrelevant to the current task or contains repetitive technical noise.
  • Archive: When information must be preserved (e.g., decisions, ADRs) but is not needed for the immediate turn.

7. Measuring Cache Hit Rate

Continuous improvement relies on telemetry.

  • Monitoring: Track Cache Hit Rate (CHR) metrics provided by the API provider.
  • Iteration: Analyze sessions with low CHR. Are the system prompts shifting? Is tool definition order inconsistent?
  • Optimization Loop:
    1. Measure performance per turn.
    2. Identify volatility in the prefix.
    3. Refactor static content to improve cache stability.
    4. Re-measure.

Best Practices Checklist

  • Does my system prompt remain stable across requests?
  • Are tools sorted consistently?
  • Is critical information pinned to the Head or Tail?
  • Have I pruned unused code/data from the ingestion stream?
  • Is there an automated summarization step for history?

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

同じリポジトリのスキル

概要と使いどころ

2d-games

無料

2D game development principles. Sprites, tilemaps, physics, camera.

日本語の概要は準備中です。原文の説明を表示しています。

VoDaiLocz/kilo-kit-mcp272026年9月13日 更新

3d-games

無料

3D game development principles. Rendering, shaders, physics, cameras.

日本語の概要は準備中です。原文の説明を表示しています。

VoDaiLocz/kilo-kit-mcp272026年9月13日 更新

aesthetic

無料

Create aesthetically beautiful interfaces following proven design principles. Use when building UI/UX, analyzing designs from inspiration sites, generating design images with ai-multimodal, implementing visual hierarchy and color theory, adding micro-interactions, or creating design documentation. Includes workflows for capturing and analyzing inspiration screenshots with chrome-devtools and ai-multimodal, iterative design image generation until aesthetic standards are met, and comprehensive design system guidance covering BEAUTIFUL (aesthetic principles), RIGHT (functionality/accessibility), SATISFYING (micro-interactions), and PEAK (storytelling) stages. Integrates with chrome-devtools, ai-multimodal, media-processing, ui-styling, and web-frameworks skills.

日本語の概要は準備中です。原文の説明を表示しています。

VoDaiLocz/kilo-kit-mcp272026年9月13日 更新

Use when implementing or managing persistent, hierarchical memory systems for AI agents. Covers cross-session state, fact supersession, and self-managed memory tools to enable long-term recall and adaptive agent behavior.

日本語の概要は準備中です。原文の説明を表示しています。

VoDaiLocz/kilo-kit-mcp272026年9月13日 更新

Use when monitoring, tracing, or debugging agentic workflows in production. Keywords: observability, tracing, OpenTelemetry, Langfuse, latency, token cost, loop detection, telemetry.

日本語の概要は準備中です。原文の説明を表示しています。

VoDaiLocz/kilo-kit-mcp272026年9月13日 更新

Use when building self-correcting retrieval systems for AI agents. Keywords: RAG, retrieval, Corrective RAG, Self-RAG, query decomposition, reranking, hallucination, grounding.

日本語の概要は準備中です。原文の説明を表示しています。

VoDaiLocz/kilo-kit-mcp272026年9月13日 更新

VoDaiLocz のスキルをすべて見る

このスキルの問題を報告する