本文へ移動
cccskills
無料GitHub で公開

skillforge

Use when creating, improving, finding, or auditing agent skills - the user says 'create a skill', 'do I have a skill for X', 'improve the X skill', 'which skill should I use', asks whether a skill exists for a task, or wants to validate, test, evaluate, package, or health-check skills. Also use for skill ecosystem maintenance (duplicate detection, stale skills, trigger collisions) and advisor checkpoints.

インストール方法を見る

含まれるファイル(76)

  • SKILL.md8.9 KB
  • .gitignore541 B
  • .skillignore289 B
  • assets/images/01-title.png1.2 MB
  • assets/images/02-quality-gap.png1.2 MB
  • assets/images/03-quality-built-in.png1.3 MB
  • assets/images/04-four-phase-architecture.png973.1 KB
  • assets/images/05-phase1-thinking-lenses.png826.6 KB
  • assets/images/06-phases-2-3.png789.1 KB
  • assets/images/07-phase4-synthesis.png752.9 KB
  • assets/images/08-evolution-mandate.png863.1 KB
  • assets/images/09-core-principles.png900.2 KB
  • assets/images/10-agentic-capabilities.png1.0 MB
  • assets/images/11-directory-structure.png995.5 KB
  • assets/images/12-installation.png839.7 KB
  • assets/images/13-closing.png1.2 MB
  • assets/templates/github-workflow-skill-ci.yml4.3 KB
  • assets/templates/script-template.py8.6 KB
  • assets/templates/skill-md-template.md1.4 KB
  • commands/skillforge.md1.4 KB
  • CONTEXT.md8.6 KB
  • docs/adr/0001-proactive-context-skill-advisor.md2.2 KB
  • docs/adr/0002-hooks-based-advisor.md5.6 KB
  • index.html69.6 KB
  • LICENSE1.0 KB
  • README.md5.2 KB
  • references/claude-code-frontmatter.md3.2 KB
  • references/degrees-of-freedom.md4.2 KB
  • references/evolution-scoring.md8.8 KB
  • references/iteration-guide.md3.7 KB
  • references/multi-lens-framework.md10.3 KB
  • references/regression-questions.md11.1 KB
  • references/script-integration-framework.md17.0 KB
  • references/script-patterns-catalog.md21.0 KB
  • references/specification-template.md2.7 KB
  • references/synthesis-protocol.md3.6 KB
  • references/testing-and-evals.md5.5 KB
  • scripts/_constants.py8.3 KB
  • scripts/advisor_scoring.py17.4 KB
  • scripts/check_docs_safety.py2.5 KB
  • scripts/common.py1.7 KB
  • scripts/compile_skill.py11.3 KB
  • scripts/context_advisor.py13.2 KB
  • scripts/context_sources.py11.3 KB
  • scripts/discover_skills.py14.8 KB
  • scripts/frontmatter.py15.4 KB
  • scripts/hooks/session_start.py4.4 KB
  • scripts/hooks/user_prompt_submit.py6.3 KB
  • scripts/init_skill.py12.5 KB
  • scripts/install_skillforge.py14.3 KB
  • scripts/install_workshop.sh5.9 KB
  • scripts/mine_skill_friction.py15.7 KB
  • scripts/package_skill.py6.9 KB
  • scripts/quick_validate.py3.2 KB
  • scripts/run_skill_evals.py24.8 KB
  • scripts/skillforge_config.py7.2 KB
  • scripts/skillforge_doctor.py16.4 KB
  • scripts/tests/fixtures/sample-skill/SKILL.md1.4 KB
  • scripts/tests/test_compile.py9.6 KB
  • scripts/tests/test_config_installer.py7.4 KB
  • scripts/tests/test_config_privacy.py7.1 KB
  • scripts/tests/test_context_advisor.py19.4 KB
  • scripts/tests/test_discovery.py8.5 KB
  • scripts/tests/test_doctor.py11.8 KB
  • scripts/tests/test_friction.py12.8 KB
  • scripts/tests/test_frontmatter.py11.4 KB
  • scripts/tests/test_hooks.py12.7 KB
  • scripts/tests/test_init_skill.py5.2 KB
  • scripts/tests/test_package_skill_ignore.py4.3 KB
  • scripts/tests/test_run_skill_evals.py16.7 KB
  • scripts/tests/test_triage.py8.4 KB
  • scripts/tests/test_validate_skill.py10.0 KB
  • scripts/triage_skill_request.py28.0 KB
  • scripts/validate_skill.py30.7 KB
  • scripts/validate-skill.py657 B
  • SKILLFORGE_AUDIT.md21.9 KB

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

SkillForge 6 - Skill Router, Creator & Ecosystem Maintainer

Routes any skill-related request to the right action (use, improve, create, compose), creates new skills through an evidence-driven pipeline, and maintains the health of the whole skill ecosystem. Core principle: skill quality is a property of behavior, not documents - a skill is done when a fresh agent demonstrably does better with it than without it.

Routing (Phase 0)

Always triage before creating anything:

python3 scripts/discover_skills.py            # refresh index (auto-refreshes if >24h old)
python3 scripts/triage_skill_request.py "<the user's request>" --json
Triage resultAction
Strong match (existing skill)Recommend it; do not create a duplicate
Moderate matchOffer IMPROVE_EXISTING on the matched skill
Weak/no match + create intentProceed to creation pipeline
Multi-domainSuggest composing existing skills
AmbiguousAsk one clarifying question

Match bands are keyword-evidence heuristics, not calibrated probabilities - report them as "strong/moderate/weak match", never as percent confidence.

Creation pipeline

Run phases in order. Each phase's detailed procedure lives in its reference - read the reference when you reach the phase, not before.

0. Baseline gate (RED). Before designing anything, dispatch a fresh subagent (Task tool) on 1-2 representative target tasks WITHOUT the skill. Capture verbatim what it does wrong. If the baseline does not fail, stop - the skill is unnecessary. The failures become the skill's test cases and its description keywords. See references/testing-and-evals.md.

1. Analysis. Identify explicit, implicit, and discovered requirements. Apply the three load-bearing lenses - Inversion (what guarantees failure → anti-patterns), Pareto (which 20% of scope delivers 80% → cut the rest), Root Cause (is this the real problem?) - plus any others from references/multi-lens-framework.md that earn their tokens. Classify the failure type you are guarding against and match the guidance form to it (see the failure-form table in references/testing-and-evals.md). Choose instruction specificity with references/degrees-of-freedom.md. Decide scripts with references/script-integration-framework.md.

2. Specification. Write the spec using references/specification-template.md. Minimal tier (problem, requirements, decisions with WHY, success criteria, test scenarios) for most skills; full tier (temporal projection, obsolescence triggers, extension points) only for infrastructure skills. Never fill a section you cannot ground - omit it.

3. Generation in fresh context. Dispatch a subagent (Task tool) that receives ONLY the spec and the baseline failures - not the analysis transcript - to write SKILL.md and supporting files. Scaffold first: python3 scripts/init_skill.py <name> --path <skills-dir>. Description doctrine: trigger conditions only, third person, symptom keywords, never a workflow summary. Budget: SKILL.md under 1,500 words; move depth to references/; <details> tags save zero tokens for agents - do not use them.

4. Execution testing (GREEN). Re-run the baseline tasks WITH the skill via fresh subagents. Gate on behavioral delta: the with-skill runs must not exhibit the baseline failures. Then run the description-triggering check (positive and near-miss queries). Iterate description and body against observed failures, not hunches. For improvements to existing skills, use blind A/B judging. Full protocols: references/testing-and-evals.md.

5. Review = lint + one adversarial reviewer. Mechanical gates first:

python3 scripts/validate_skill.py <skill-dir>     # structure, frontmatter, lint (pinned models, word budget, description shape)
python3 scripts/check_docs_safety.py <skill-dir>

Then one fresh-context subagent prompted to REFUTE the skill (find the case where it misleads, over-triggers, or fails its own scenarios), carrying the reviewer checklists in references/synthesis-protocol.md. Fix what it proves; ship what survives. Do not convene approval panels - same-model unanimity measures nothing.

6. Ship with evals. Every generated skill keeps its tests: an evals/ directory (trigger queries + behavioral scenarios + assertions) so future edits can be regression-tested with python3 scripts/run_skill_evals.py <skill-dir>. Iterate post-ship with references/iteration-guide.md.

Frontmatter and platform facts

Write frontmatter against the current Claude Code field set (17 fields) documented in references/claude-code-frontmatter.md, which also covers hooks (hooks receive JSON on stdin, not env vars), context: fork/agent, $ARGUMENTS, and the agentskills.io portability limits (64-char name, 1024-char description) that validate_skill.py enforces. Never pin dated model IDs (claude-*-YYYYMMDD) - the validator rejects them.

Ecosystem maintenance

python3 scripts/skillforge_doctor.py              # trigger collisions, duplicates, stale refs, token budgets, description lint
python3 scripts/compile_skill.py <dir> --target claude|codex|agentskills
python3 scripts/package_skill.py <dir> ./dist     # .skill zip, honors .skillignore
python3 scripts/mine_skill_friction.py --consent  # opt-in: mine local transcripts for skill friction

Use doctor output to drive IMPROVE_EXISTING work; use friction reports as advisor evidence.

Context Skill Advisor

Proactive suggestions are delivered through Claude Code hooks (SessionStart surfaces the queue; UserPromptSubmit scores checkpoints inline) - no daemon. Configure with python3 scripts/install_skillforge.py (interactive; hooks and Personal Context scanning are opt-in, never default). Manage the queue: python3 scripts/context_advisor.py list|use|snooze|dismiss. Suggestions are evidence-backed and never auto-invoke a skill.

Script inventory

ScriptPurpose
discover_skills.pyBuild/refresh the cross-runtime skill index
triage_skill_request.pyRoute input to use/improve/create/compose/clarify
validate_skill.pyFull structural + lint validation (quick_validate.py = fast subset)
run_skill_evals.pyRun a skill's evals/ regression suite
skillforge_doctor.pyEcosystem health report
init_skill.pyScaffold a new skill (with evals/)
compile_skill.pyCompile a skill for a target runtime
package_skill.pyPackage as .skill archive
mine_skill_friction.pyOpt-in transcript friction mining
context_advisor.py / install_skillforge.pyAdvisor queue and setup
check_docs_safety.pyUnsafe interpolation check

Script exit codes: 0 success, 1 failure, 2 usage/consent error, 10 validation failure, 11 verification/dependency failure.

Extension points: new lint checks in validate_skill.py; new doctor checks in skillforge_doctor.py; new compile targets in compile_skill.py; new lenses in references/multi-lens-framework.md.

Anti-patterns

AvoidInstead
Creating without a failing baselineRun the RED gate; no failure = no skill
Description that summarizes workflowTrigger conditions only - agents act on summaries and skip the body
Body "Triggers" sections as a mechanismOnly the frontmatter description drives invocation
Approval panels and self-scored gatesLint what is falsifiable; adversarially refute the rest
<details> blocks for "progressive disclosure"Separate reference files loaded on demand
Pinned dated model IDsFamily aliases or omit model:
Duplicating an existing skillPhase 0 triage first, always

Verification checklist

  • Baseline failure captured before writing (RED)
  • With-skill runs clear the baseline failures (GREEN)
  • Trigger check passes on positive and near-miss queries
  • validate_skill.py and check_docs_safety.py pass
  • Adversarial reviewer's proven issues fixed
  • evals/ shipped with the skill; run_skill_evals.py passes
  • SKILL.md under 1,500 words (wc -w)

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

このスキルの問題を報告する