本文へ移動
cccskills
無料GitHub で公開

toolkit

Toolkit management: create and evaluate skills and agents, manage routing tables, generate Claude.md.

インストール方法を見る

含まれるファイル(79)

  • SKILL.md12.0 KB
  • agents/skill-creator/analyzer.md4.8 KB
  • agents/skill-creator/comparator.md4.7 KB
  • agents/skill-creator/grader.md3.9 KB
  • assets/skill-creator/eval_viewer.html57.1 KB
  • references/agent-comparison.md11.5 KB
  • references/agent-comparison/benchmark-tasks.md6.8 KB
  • references/agent-comparison/do-creation-compliance-tasks.json8.5 KB
  • references/agent-comparison/examples-and-errors.md1.0 KB
  • references/agent-comparison/grading-rubric.md4.3 KB
  • references/agent-comparison/methodology.md7.4 KB
  • references/agent-comparison/optimization-guide.md11.9 KB
  • references/agent-comparison/optimization-tasks.example.json853 B
  • references/agent-comparison/optimize-phase.md9.2 KB
  • references/agent-comparison/read-only-ops-short-tasks.json375 B
  • references/agent-comparison/report-template.md3.8 KB
  • references/agent-comparison/socratic-debugging-body-short-tasks.json372 B
  • references/agent-comparison/socratic-debugging-trigger-tasks.json2.8 KB
  • references/agent-creator.md8.8 KB
  • references/agent-creator/agent-design-patterns.md12.0 KB
  • references/agent-creator/agent-eval-design.md7.3 KB
  • references/agent-creator/agent-frontmatter-template.md8.2 KB
  • references/agent-evaluation.md10.6 KB
  • references/agent-evaluation/batch-evaluation.md1.6 KB
  • references/agent-evaluation/common-issues.md2.1 KB
  • references/agent-evaluation/report-templates.md2.2 KB
  • references/agent-evaluation/scoring-rubric.md2.4 KB
  • references/generate-claudemd.md11.0 KB
  • references/generate-claudemd/CLAUDEMD_TEMPLATE.md4.1 KB
  • references/generate-claudemd/examples-and-errors.md7.4 KB
  • references/routing-table-updater.md10.1 KB
  • references/routing-table-updater/batch-mode.md1.9 KB
  • references/routing-table-updater/conflict-resolution.md3.7 KB
  • references/routing-table-updater/error-handling.md1.9 KB
  • references/routing-table-updater/examples.md4.7 KB
  • references/routing-table-updater/extraction-patterns.md2.2 KB
  • references/routing-table-updater/routing-format.md3.4 KB
  • references/routing-table-updater/skill-examples.md2.7 KB
  • references/skill-composer.md9.1 KB
  • references/skill-composer/compatibility-matrix.md13.0 KB
  • references/skill-composer/composition-patterns.md4.9 KB
  • references/skill-composer/examples.md17.0 KB
  • references/skill-composer/skill-patterns.md12.5 KB
  • references/skill-creator.md28.9 KB
  • references/skill-creator/agent-template.md17.9 KB
  • references/skill-creator/artifact-schemas.md8.2 KB
  • references/skill-creator/bundled-components.md2.3 KB
  • references/skill-creator/complexity-tiers.md8.0 KB
  • references/skill-creator/domain-research-targets.md13.2 KB
  • references/skill-creator/enrichment-workflow.md10.1 KB
  • references/skill-creator/error-catalog.md9.9 KB
  • references/skill-creator/preferred-patterns.md11.7 KB
  • references/skill-creator/progressive-disclosure.md8.5 KB
  • references/skill-creator/skill-template.md14.4 KB
  • references/skill-creator/workflow-patterns.md7.6 KB
  • references/toolkit-evolution.md14.5 KB
  • references/toolkit-evolution/diagnose-scripts.md9.5 KB
  • references/toolkit-evolution/evolution-history.md5.5 KB
  • references/toolkit-evolution/evolution-report-template.md716 B
  • references/toolkit-evolution/evolve-preferred-patterns.md6.9 KB
  • references/toolkit-evolution/evolve-scripts.md4.2 KB
  • references/weak-model-uplift.md10.0 KB
  • scripts/agent-comparison/compare.py2.3 KB
  • scripts/agent-comparison/generate_variant.py19.5 KB
  • scripts/agent-comparison/optimize_loop.py89.3 KB
  • scripts/agent-comparison/trigger_eval.py19.5 KB
  • scripts/routing-table-updater/extract_metadata.py10.8 KB
  • scripts/routing-table-updater/generate_routes.py10.5 KB
  • scripts/routing-table-updater/scan.py4.7 KB
  • scripts/routing-table-updater/update_routing.py12.4 KB
  • scripts/routing-table-updater/validate.py7.3 KB
  • scripts/skill-composer/build_dag.py12.5 KB
  • scripts/skill-composer/discover_skills.py12.1 KB
  • scripts/skill-composer/validate.py12.3 KB
  • scripts/skill-creator/aggregate_benchmark.py10.0 KB
  • scripts/skill-creator/eval_compare.py10.2 KB
  • scripts/skill-creator/optimize_description.py12.5 KB
  • scripts/skill-creator/package_results.py7.8 KB
  • scripts/skill-creator/run_eval.py6.3 KB

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

Toolkit

Nine modes covering the full toolkit lifecycle: creating and improving skills and agents; evaluating agents; maintaining routing tables; generating CLAUDE.md; composing multi-skill DAGs; and running the evolution loop. Classify the request and follow the matching section.

Mode Selection

ModeSignalsSection
Skill Creatorcreate skill, scaffold skill, new skill, build a skillCreate Skill
Agent Creatorcreate agent, scaffold agent, new agentCreate Agent
Weak-Model Upliftweaker model, uplift skill, make a skill work for Opus 4.6, improve guidance from generated outputUplift for Weaker Models
Agent Comparisoncompare agents, A/B test agents, benchmark agents, benchmark skill, bake-offCompare Agents
Agent Evaluationevaluate agent quality, audit agent, grade agent, eval skillEvaluate Agent
Skill Composercompose skills, DAG orchestration, skill pipelineCompose Skills
Routing Tablesupdate routing tables, sync routing, routing driftUpdate Routing
Toolkit Evolutionevolve toolkit, self-improve, discover gapsEvolve Toolkit
Generate CLAUDE.mdgenerate claude.md, create claude.md, initGenerate CLAUDE.md

Create Skill

Phases: INTENT -> DRAFT -> TEST -> REGISTER

  1. Capture intent. What should the skill do? When should it trigger? What output? Are outputs objectively verifiable (code, data) or subjective (writing, design)?
  2. Duplicate check. Run grep -i "<domain>" skills/*/SKILL.md to check existing coverage. If an umbrella skill covers the domain, add a reference file instead.
  3. Write SKILL.md. Follow references/skill-creator/skill-template.md for frontmatter structure. Apply Dense-Complete Writing standard. Frontmatter must include: name, description, routing (triggers, not_for, category, pairs_with), allowed-tools.
  4. Test. Try 3 should-trigger, 2 should-not-trigger, and 2 near-miss prompts with the skill loaded. Revise the SKILL.md until routing and output are right.
  5. Register. Run python3 scripts/generate-skill-index.py to update routing.

Load references/skill-creator.md for the full workflow. Deep references in references/skill-creator/ cover progressive disclosure, artifact schemas, complexity tiers, error catalog, enrichment workflow, and more.

Scripts: scripts/skill-creator/


Create Agent

Phases: DISCOVER -> DESIGN -> SCAFFOLD -> REGISTER -> VALIDATE

  1. Discover. Check for domain overlap: grep -i "<domain>" agents/*.md. If an existing agent covers the domain, add a references/ file instead.
  2. Design. Decide role type (reviewer/engineer/orchestrator), allowed tools, complexity, triggers (3-6 specific phrases), pairs_with (verify each exists), reference files, description (intent verb + domain + boundary clause), activation cases.
  3. Scaffold. Write the agent file using references/agent-creator/agent-frontmatter-template.md. Follow docs/PHILOSOPHY.md for operator context structure.
  4. Register. Run python3 scripts/generate-agent-index.py.
  5. Validate. Run python3 scripts/validate-references.py to check reference file integrity. Test activation with the 3+2+2 prompt set.

Load references/agent-creator.md for full phases. Deep references in references/agent-creator/ cover design patterns, frontmatter template, eval design.


Uplift for Weaker Models

Improve a skill, agent, or shared guide until a weaker model produces strong output with it. Load references/weak-model-uplift.md and follow its steps:

  1. Pick the target from data. Query ~/.claude/learning/usage.db and learning.db for heavily used or failing skills.
  2. Build tasks and checks first. 4–8 tasks plus 1–2 held-out tasks; deterministic checks and a yes/no rubric written before any run.
  3. Run the arms. No guidance and current guidance, two samples per task minimum, with python3 scripts/weak_model_run.py.
  4. Score and look. Checks, rubric, your own review of every artifact, optional Jev questions on extracted facts.
  5. Turn failures into rules. Concrete values, before/after examples, runnable checks; delete stale instructions; examples from unrelated products.
  6. Rerun the guided arm and held-out tasks; stop when gains flatten or after three rounds.
  7. Report and ship a per-round table with held-out results, cost, and caveats in the PR body.

Compare Agents

Controlled benchmarks comparing agent variants on identical tasks.

  1. Select variants. Identify the agents to compare (2-4 variants).
  2. Design benchmark. Load references/agent-comparison/benchmark-tasks.md. Select 5-10 representative tasks covering the agent's domain.
  3. Execute. Run each task with each variant. Collect: output quality, token usage, tool calls, time.
  4. Grade. Apply rubric from references/agent-comparison/grading-rubric.md. Score each dimension.
  5. Report. Use references/agent-comparison/report-template.md. Include: methodology, per-task scores, aggregate rankings, cost analysis, recommendation.
  6. Optimize. Load references/agent-comparison/optimize-phase.md to improve the winning variant further.

Load references/agent-comparison.md for the full methodology.


Evaluate Agent

Static structural and standards-compliance grading with a 90-point deterministic scorer.

  1. Read the agent file. Extract frontmatter, body sections, reference files.
  2. Score. Apply rubric from references/agent-evaluation/scoring-rubric.md. Categories: identity (15 pts), expertise (20 pts), routing (15 pts), references (15 pts), workflow (15 pts), standards (10 pts).
  3. Report. Use references/agent-evaluation/report-templates.md. Include: per-category scores, specific findings, improvement recommendations.
  4. Batch mode. For multiple agents: references/agent-evaluation/batch-evaluation.md.

Load references/agent-evaluation.md for the full methodology.


Compose Skills

DAG-based multi-skill orchestration with dependency resolution.

  1. Define the DAG. List skills in execution order. Identify dependencies (skill B needs output from skill A).
  2. Check compatibility. Load references/skill-composer/compatibility-matrix.md. Verify input/output contracts between skills.
  3. Build the pipeline. Load references/skill-composer/composition-patterns.md for orchestration patterns (serial, parallel, fan-out, conditional).
  4. Execute. Run skills in DAG order. Pass outputs between skills via the defined contracts.
  5. Validate. Check all skills completed. Verify final output meets the composite goal.

Load references/skill-composer.md for the full methodology. See references/skill-composer/examples.md for worked examples.

Scripts: scripts/skill-composer/


Update Routing

5-phase pipeline: SCAN -> EXTRACT -> GENERATE -> UPDATE -> VERIFY.

  1. SCAN. Run python3 scripts/generate-skill-index.py to discover all skills and agents.
  2. EXTRACT. Parse frontmatter from each SKILL.md and agent file. Extract triggers, description, category, complexity.
  3. GENERATE. Build skills/INDEX.json and agents/INDEX.json.
  4. UPDATE. Write index files. PostToolUse hooks auto-regenerate on individual edits; this covers bulk changes and drift.
  5. VERIFY. Compare generated index against discovered files. Report missing entries, conflicts, or stale entries.

Load references/routing-table-updater.md for full phases. Deep references in references/routing-table-updater/ cover routing format, extraction patterns, conflict resolution, batch mode.


Evolve Toolkit

7-phase pipeline: DISCOVER -> DIAGNOSE -> PROPOSE -> CRITIQUE -> BUILD -> VALIDATE -> EVOLVE.

  1. DISCOVER. Audit recent sessions for routing failures, skill gaps, agent weaknesses, user friction.
  2. DIAGNOSE. Load references/toolkit-evolution/diagnose-scripts.md. Run gap analysis scripts. Identify patterns.
  3. PROPOSE. Generate 3-5 improvement proposals with expected impact, effort, risk.
  4. CRITIQUE. Apply multi-perspective review to proposals.
  5. BUILD. Implement the approved proposals using the appropriate mode above (create skill, create agent, etc.).
  6. VALIDATE. Run tests and validators on new/changed components.
  7. EVOLVE. Update evolution history at references/toolkit-evolution/evolution-history.md.

Load references/toolkit-evolution.md for the full pipeline.


Generate CLAUDE.md

4-phase pipeline: SCAN -> DETECT -> GENERATE -> VALIDATE.

  1. SCAN. Check for existing CLAUDE.md. If present, write to CLAUDE.md.generated for comparison. Detect language, framework, build system from repo files.
  2. DETECT. Identify domain enrichment opportunities. Load references/generate-claudemd/examples-and-errors.md for language-specific patterns.
  3. GENERATE. Load template from references/generate-claudemd/CLAUDEMD_TEMPLATE.md. Fill sections: overview, commands, architecture, conventions, testing, deployment.
  4. VALIDATE. Run all documented commands. Verify paths exist. Check for secrets in output.

Optional modes: subdirectory CLAUDE.md for monorepos; minimal mode (overview + commands + architecture only).


Deep References

Load when the task needs detailed schemas, templates, or methodology.

ModeKey References
Skill Creatorreferences/skill-creator.md, references/skill-creator/{skill-template,progressive-disclosure,complexity-tiers,error-catalog,enrichment-workflow}.md
Agent Creatorreferences/agent-creator.md, references/agent-creator/{agent-design-patterns,agent-frontmatter-template,agent-eval-design}.md
Weak-Model Upliftreferences/weak-model-uplift.md
Agent Comparisonreferences/agent-comparison.md, references/agent-comparison/{methodology,grading-rubric,benchmark-tasks,report-template,optimize-phase}.md
Agent Evaluationreferences/agent-evaluation.md, references/agent-evaluation/{scoring-rubric,report-templates,batch-evaluation}.md
Skill Composerreferences/skill-composer.md, references/skill-composer/{compatibility-matrix,composition-patterns,skill-patterns,examples}.md
Routing Tablesreferences/routing-table-updater.md, references/routing-table-updater/{routing-format,extraction-patterns,conflict-resolution,examples}.md
Toolkit Evolutionreferences/toolkit-evolution.md, references/toolkit-evolution/{diagnose-scripts,evolution-history,evolve-preferred-patterns}.md
Generate CLAUDE.mdreferences/generate-claudemd.md, references/generate-claudemd/{CLAUDEMD_TEMPLATE,examples-and-errors}.md

Scripts and Agents

ModeScriptsAgents
Skill Creatorscripts/skill-creator/agents/skill-creator/
Skill Composerscripts/skill-composer/--
Weak-Model Upliftscripts/weak_model_run.py (repo root)--
Routing Tablesscripts/routing-table-updater/--
Agent Comparisonscripts/agent-comparison/--

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

同じリポジトリのスキル

概要と使いどころ

accounts

無料

Manage multiple Claude Code accounts: add, list, check, launch, and install shell aliases for 10+ isolated CLAUDE_CONFIG_DIR profiles.

日本語の概要は準備中です。原文の説明を表示しています。

notque/vexjoy-agent4412026年10月11日 更新

Improve architecture across modules by deepening interfaces.

日本語の概要は準備中です。原文の説明を表示しています。

notque/vexjoy-agent4412026年10月11日 更新

Assessment: read-only inspection, codebase overview, value analysis, health checks, ADR consultation, decision analysis, multi-perspective critique.

日本語の概要は準備中です。原文の説明を表示しています。

notque/vexjoy-agent4412026年10月11日 更新

Background memory consolidation — overnight review, merge, and injection payload for memory files.

日本語の概要は準備中です。原文の説明を表示しています。

notque/vexjoy-agent4412026年10月11日 更新

Jev-driven browser automation: Jev picks operations, programs execute, a text model writes field values only when Jev cannot pick one from the goal.

日本語の概要は準備中です。原文の説明を表示しています。

notque/vexjoy-agent4412026年10月11日 更新

Write, compose, integrate, and improve programs that call Jev, TypeSafe's System One judgment model.

日本語の概要は準備中です。原文の説明を表示しています。

notque/vexjoy-agent4412026年10月11日 更新

notque のスキルをすべて見る

このスキルの問題を報告する