Web・iOS・Androidの画面を、読み上げやキーボード操作に対応させ、ラベル、配色、操作対象の大きさなどをWCAG 2.2に沿って設計・点検するスキル。
- アイコンボタンの説明を付けたいとき
- キーボード操作とモーダルの点検
- コントラストや操作対象の大きさの確認
AIによる開発の成功条件を実装前に定義し、コード・ルール・モデル・人の評価で新機能と既存機能を確認しながら、複数回の試行から信頼性を記録するスキル。
原文Eval-driven development (EDD) framework for AI coding sessions — define capability and regression evals before coding, grade with code-based, model-based, rule, or human graders, and track pass@k and pass^k reliability. Use when defining pass/fail criteria for agent tasks, measuring agent reliability, building regression suites for prompt or agent changes, or benchmarking across model versions.
インストール方法を見るAIによる開発で、実装前に成功条件を定め、実装中も評価を繰り返す手順を整えます。新たな作業能力の評価と、変更後も既存機能が動くかを確かめる回帰評価を用意し、コード、ルール、モデル、人による採点を使い分けます。複数回のうち一度でも成功するpass@kと、全試行が成功するpass^kを記録し、結果を報告します。
エージェントの完了条件を明確にしたいときや、プロンプト・エージェントの変更による品質低下を調べたいときに向いています。モデルのバージョンごとの性能比較や、重要な処理が繰り返し安定して成功するかの確認にも使えます。
Claude Code向けの手順と、Node.jsで使うローカルユーティリティが含まれます。ユーティリティの候補コード実行は、検証済みのOS隔離基盤がないため全OSで無効です。静的警告や記録の検証結果は、安全な実行や採用の証拠にはなりません。セキュリティ確認には人のレビューを求めています。
この紹介文は、公開されている SKILL.md をもとに AI(Claude Haiku)が作成しました。正確な仕様は下の原文を確認してください。
インストールする前に、エージェントに与えられる指示の中身を確認できます。
A formal evaluation framework for Claude Code sessions, implementing eval-driven development (EDD) principles.
Eval-Driven Development treats evals as the "unit tests of AI development":
Test if Claude can do something it couldn't before:
[CAPABILITY EVAL: feature-name]
Task: Description of what Claude should accomplish
Success Criteria:
- [ ] Criterion 1
- [ ] Criterion 2
- [ ] Criterion 3
Expected Output: Description of expected result
Ensure changes don't break existing functionality:
[REGRESSION EVAL: feature-name]
Baseline: SHA or checkpoint name
Tests:
- existing-test-1: PASS/FAIL
- existing-test-2: PASS/FAIL
- existing-test-3: PASS/FAIL
Result: X/Y passed (previously Y/Y)
Deterministic checks using code:
# Check if file contains expected pattern
grep -q "export function handleAuth" src/auth.ts && echo "PASS" || echo "FAIL"
# Check if tests pass
npm test -- --testPathPattern="auth" && echo "PASS" || echo "FAIL"
# Check if build succeeds
npm run build && echo "PASS" || echo "FAIL"
Use Claude to evaluate open-ended outputs:
[MODEL GRADER PROMPT]
Evaluate the following code change:
1. Does it solve the stated problem?
2. Is it well-structured?
3. Are edge cases handled?
4. Is error handling appropriate?
Score: 1-5 (1=poor, 5=excellent)
Reasoning: [explanation]
Flag for manual review:
[HUMAN REVIEW REQUIRED]
Change: Description of what changed
Reason: Why human review is needed
Risk Level: LOW/MEDIUM/HIGH
"At least one success in k attempts"
"All k trials succeed"
## EVAL DEFINITION: feature-xyz
### Capability Evals
1. Can create new user account
2. Can validate email format
3. Can hash password securely
### Regression Evals
1. Existing login still works
2. Session management unchanged
3. Logout flow intact
### Success Metrics
- pass@3 > 90% for capability evals
- pass^3 = 100% for regression evals
Write code to pass the defined evals.
# Run capability evals
[Run each capability eval, record PASS/FAIL]
# Run regression evals
npm test -- --testPathPattern="existing"
# Generate report
EVAL REPORT: feature-xyz
========================
Capability Evals:
create-user: PASS (pass@1)
validate-email: PASS (pass@2)
hash-password: PASS (pass@1)
Overall: 3/3 passed
Regression Evals:
login-flow: PASS
session-mgmt: PASS
logout-flow: PASS
Overall: 3/3 passed
Metrics:
pass@1: 67% (2/3)
pass@3: 100% (3/3)
Status: READY FOR REVIEW
/eval define feature-name
Creates eval definition file at .claude/evals/feature-name.md
/eval check feature-name
Runs current evals and reports status
/eval report feature-name
Generates full eval report
Store evals in project:
.claude/
evals/
feature-xyz.md # Eval definition
feature-xyz.log # Eval run history
baseline.json # Regression baselines
## EVAL: add-authentication
### Phase 1: Define (10 min)
Capability Evals:
- [ ] User can register with email/password
- [ ] User can login with valid credentials
- [ ] Invalid credentials rejected with proper error
- [ ] Sessions persist across page reloads
- [ ] Logout clears session
Regression Evals:
- [ ] Public routes still accessible
- [ ] API responses unchanged
- [ ] Database schema compatible
### Phase 2: Implement (varies)
[Write code]
### Phase 3: Evaluate
Run: /eval check add-authentication
### Phase 4: Report
EVAL REPORT: add-authentication
==============================
Capability: 5/5 passed (pass@3: 100%)
Regression: 3/3 passed (pass^3: 100%)
Status: SHIP IT
The mechanical utilities ship in scripts/lib/eval-harness/:
node scripts/eval-harness.js example
node scripts/eval-harness.js capsule group <dir> [<dir> ...]
groups 1 to 100 explicitly selected, verified local capsule snapshots from one
task family by declared harness version. Repeated snapshots count once;
conflicting identities or invalid capsules reject the whole report. This is
read-only record counting, with no new rollouts, scores or promotion. Use small,
quiescent capsules. Payloads, directory arguments and raw run/capsule IDs are
omitted, but task-family/version labels are verbatim and digest references are
linkable; review them before sharing. Operational validation remains pending.Candidate execution is disabled on every OS because no verified OS containment
backend is implemented. gate run, runGate, runVariant, direct child launch,
and the retired effect preload refuse with gate.isolation_required. No trust
flag or caller-supplied executor can bypass the refusal. The example records
that refusal and inspects source without executing or scoring it.
Do not present static warnings, a capsule receipt, or successful utility tests
as candidate containment or promotion evidence. A future gate requires an
independently reviewed OS boundary, protected checker and audit channels, and
fatal baseline rejection. See docs/architecture/eval-harness-frameworks.md.
Use product evals when behavior quality cannot be captured by unit tests alone.
pass@1: direct reliabilitypass@3: practical reliability under controlled retriespass^3: stability test (all 3 runs must pass)Recommended thresholds:
.claude/evals/<feature>.md definition.claude/evals/<feature>.log run historydocs/releases/<version>/eval-summary.md release snapshotまだレビューはありません。使ってみた感想をお寄せください。
概要と使いどころ
Web・iOS・Androidの画面を、読み上げやキーボード操作に対応させ、ラベル、配色、操作対象の大きさなどをWCAG 2.2に沿って設計・点検するスキル。
AIエージェントの不調を、指示・記憶・ツール実行・画面表示など12の層から調べるスキル。コードやログを根拠に原因を整理し、重要度順の指摘と修正案をまとめます。
実際の開発課題で複数のコーディングエージェントを比較するスキル。成功率、取得可能なAPI費用、所要時間、繰り返し実行の安定性を測り、選定や更新後の評価に使えます。
AIエージェントが使うツールの種類や入出力、エラーからの復帰手順を設計・見直します。文脈の情報量も整理し、作業完了率や再試行回数で改善を評価します。
AIエージェントが失敗や同じ操作を繰り返す原因を、エラーと実行状況から整理します。小さな復旧操作を試し、結果と根拠を引き継げる報告にまとめるスキルです。
AIエージェントの失敗や同じ操作の繰り返しを記録し、原因の切り分け、小さな復旧操作、結果の報告まで進める手順を示して、根拠のある再試行につなげるスキル。