本文へ移動
cccskills
無料GitHub で公開

eval-runner

Run eval scenarios to benchmark Mycelium effectiveness. Execute tasks using reflexion loop, validate against success criteria, record metrics.

インストール方法を見る

含まれるファイル(1)

  • SKILL.md4.2 KB

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

Eval Runner

Benchmark the agent's performance against defined scenarios. Adapted from n-trax eval system.

Commands

run <category/name>

  1. Read YAML from .claude/evals/scenarios/<category>/<name>.yml
  2. Parse fields (name, category, task_prompt, success_criteria, budget)
  3. Execute setup steps if defined
  4. Record start time
  5. Execute task via reflexion workflow (read corrections first)
  6. Record end time and iteration count
  7. Validate ALL success criteria
  8. Write result JSON to .claude/evals/results/<timestamp>-<name>.json
  9. Report summary

run-all [category]

  1. Glob .claude/evals/scenarios/**/*.yml
  2. Skip scenarios with status: retired
  3. For each: run in isolation (git stash), record result, restore
  4. Update .claude/evals/pass-history.json with each result
  5. Aggregate and report

run-split <optimization|holdout>

  1. Glob .claude/evals/scenarios/**/*.yml
  2. Read each YAML, filter by split field matching the requested set
  3. Skip scenarios with status: retired
  4. For each matching scenario: run in isolation, record result, restore
  5. Update .claude/evals/pass-history.json with each result
  6. Aggregate and report (label output clearly as "Optimization Set" or "Holdout Set")

report

  1. Read all results from .claude/evals/results/
  2. Generate summary table:
| Category    | Pass Rate | Avg Iterations | Avg Time | Notes |
|-------------|-----------|----------------|----------|-------|
| discovery   | ...       | ...            | ...      |       |
| delivery    | ...       | ...            | ...      |       |
| integration | ...       | ...            | ...      |       |
| **Overall** | ...       | ...            | ...      |       |
  1. List failure patterns and recommendations

prune

  1. Read .claude/evals/pass-history.json
  2. Flag evals where last_5 is all-pass (saturated) or all-fail (broken)
  3. Flag evals with no runs in 30+ days (stale)
  4. For saturated evals, suggest: retire or increase difficulty
  5. For broken evals, suggest: fix criteria or retire
  6. Present recommendations — do NOT auto-retire
  7. On user confirmation: set status: retired in scenario YAML, update pass-history.json, log in .claude/harness/decision-log.md

mine

Analyze audit logs to propose new eval scenarios from observed failure patterns.

  1. Read .claude/state/change-log.jsonl (last 100 entries)
  2. Read .claude/state/diamond-state-audit.jsonl (all entries)
  3. Group change-log entries by session_id, identify: a. Correction clusters: 3+ edits to same file in one session (agent struggled) b. Skill friction: edits to .claude/skills/*/SKILL.md during a session (instructions unclear) c. Missing test coverage: 5+ files changed with no test file edits
  4. Count diamond bypass entries from audit log
  5. Cross-reference with existing scenarios in pass-history.json to avoid duplicates
  6. For each pattern, output a proposed eval scenario as YAML template
  7. Tag proposals with source: trace-mining and originating session_id
  8. Do NOT auto-create files — present for human review

See .claude/evals/schema.md §Trace Mining Heuristics for pattern-to-eval mappings.

Result Format

See .claude/evals/schema.md for YAML scenario and JSON result formats.

Pass History

After writing each result JSON (step 8 of run), also update .claude/evals/pass-history.json:

  • Increment runs and passes (if passed) for the eval
  • Append true/false to last_5 (trim to keep only last 5)
  • Update last_run timestamp

Creating Scenarios

Place YAML files in .claude/evals/scenarios/<category>/. Define task_prompt, success_criteria, and budget. Set split, status, and source fields per schema.

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

同じリポジトリのスキル

概要と使いどころ

Accessibility audit, scoped to the surfaces a product actually has. Detects web / rendered_markdown / terminal / native_app / video_audio / document / headless, then applies only the criteria that bind. WCAG 2.1 AA in full for web; not at all for headless.

日本語の概要は準備中です。原文の説明を表示しています。

haabe/mycelium462026年10月10日 更新

adopt

無料

Bring Mycelium into a project that already has code. Detects that the repo predates the framework, asks before touching anything, then reads the codebase to draft what it CAN establish (delivery, solution shape) and — the actual point — names what it cannot (purpose, strategy, real user evidence). The output is a discovery backlog with a head start, never a filled canvas.

日本語の概要は準備中です。原文の説明を表示しています。

haabe/mycelium462026年10月10日 更新

Design the smallest viable test to validate or invalidate a critical assumption. Based on Torres's assumption testing framework, organized by Gilad's AFTER model (Assessment → Fact-Finding → Tests → Experiments → Release Results).

日本語の概要は準備中です。原文の説明を表示しています。

haabe/mycelium462026年10月10日 更新

Use before any research activity or significant decision. Reviews cognitive biases relevant to the current stage.

日本語の概要は準備中です。原文の説明を表示しています。

haabe/mycelium462026年10月10日 更新

Use to evaluate whether current work aligns with Better Value Sooner Safer Happier. Run at diamond completion and periodically.

日本語の概要は準備中です。原文の説明を表示しています。

haabe/mycelium462026年10月10日 更新

Lint canvas files for staleness, missing fields, inconsistent evidence types, and orphaned references. Run periodically or before major transitions.

日本語の概要は準備中です。原文の説明を表示しています。

haabe/mycelium462026年10月10日 更新

haabe のスキルをすべて見る

このスキルの問題を報告する