本文へ移動
cccskills
無料GitHub で公開

eval-harness

Build a repeatable eval loop that grades agent output with an LLM judge, so prompt/skill changes get scored against a baseline instead of eyeballed. Reuses loopkit's verifier subagent as the grader — do not build a new one.

インストール方法を見る

含まれるファイル(1)

  • SKILL.md3.5 KB

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

Eval Harness

Every prompt tweak in a long-running agent looks like an improvement in the moment. The only way to know is a graded run against fixed inputs. Loopkit already ships .claude/agents/verifier.md — that is your grader. Do not rebuild it.

The three-stage loop

inputs.jsonl  →  runner  →  outputs.jsonl  →  verifier (per row)  →  verdicts.jsonl  →  diff vs baseline

Each stage writes to disk. No stage holds the whole run in context.

Stage 1 — inputs.jsonl

One JSON object per row: {"id": "case-01", "input": "...", "expected": "..."}.

  • 20-100 cases is enough for a signal. More is nice, not required.
  • Include known-hard cases, edge cases, and a couple of trivial ones as sanity anchors.
  • Freeze the file. Rev the eval with a suffix (inputs-v2.jsonl) when you change it. Never edit in place — you lose the baseline.

Stage 2 — runner

A dumb loop: for each row, call the model with the current prompt/skill, capture output, write {"id": ..., "output": ...} to outputs.jsonl. No grading here — just capture.

  • Same temperature every run (usually 0 for evals).
  • Same seed / model version.
  • Log the git SHA of the prompt/skill under test in the file header.

If the runner is smart it will bias the eval. Keep it dumb.

Stage 3 — verifier

Fan out one subagent per row (see subagent-fanout). Each gets:

  • The input.
  • The expected output (or spec).
  • The actual output.
  • The verifier system prompt from .claude/agents/verifier.md.

Verifier returns strict JSON: {"pass": bool, "why": "..."}. Collect into verdicts.jsonl.

Diff vs baseline

Two runs of the same eval on two prompt versions → compare pass rates per case. What matters:

  • Overall pass rate — the headline.
  • Regressions — cases that were green and went red. These block ship.
  • New passes — cases that were red and went green. These justify ship.
  • Flappy cases — inconsistent across reruns. Investigate; may be genuine model nondeterminism or a bad case.

A change that raises the mean but adds regressions is usually a loss — the new failures are cases you already knew worked.

Red flags

  • Grader is the same model that produced the output, with the same prompt. Self-grading is lenient. Use a different persona at minimum; ideally a different model tier.
  • Eval passes 100% on day one. The cases are too easy, or the grader is a rubber stamp. Add adversarial cases.
  • Eval takes >30 minutes. Fan out. A serial 100-case eval is a serial 100-case bottleneck.
  • Grader sees your prompt under test. It will grade what you wanted, not what happened. Feed it only spec + input + output.
  • Baseline lost. Without baseline, "improvement" is vibes. Commit verdicts.jsonl to git.

When NOT to do this

  • One-off script — build a checklist, not a harness.
  • Prompt that changes daily and won't stabilize — evals need a fixed target.
  • Task where "correct" isn't checkable (open-ended creative writing) — use human eval or a rubric-based grader, not pass/fail.

The verifier is already yours. The harness is 100 lines of glue around it.

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

同じリポジトリのスキル

概要と使いどころ

a11y-pass

無料

Catch the accessibility failures that ship in almost every AI-built UI. Use after building any interactive component.

日本語の概要は準備中です。原文の説明を表示しています。

Archive228/loopkit7552026年7月15日 更新

Before compaction Loopkit extracts decisions into claude-decisions.json (machine-readable). Read it alongside claude-progress.txt at session start — prose is for humans, JSON is for the loop.

日本語の概要は準備中です。原文の説明を表示しています。

Archive228/loopkit7552026年7月15日 更新

Review a diff against the goal spec assuming the code is BROKEN. The reviewer that lives in the maker's head always agrees with itself — this pulls review into a hostile, separate pass. Invoke after every code change before marking work done.

日本語の概要は準備中です。原文の説明を表示しています。

Archive228/loopkit7552026年7月15日 更新

Verify that an endpoint checks ownership, not just authentication. Use on any handler that reads or mutates user data.

日本語の概要は準備中です。原文の説明を表示しています。

Archive228/loopkit7552026年7月15日 更新

Find the exact commit that introduced a bug. Use when something worked before and broke, and you don't know which change did it.

日本語の概要は準備中です。原文の説明を表示しています。

Archive228/loopkit7552026年7月15日 更新

Before picking new work, smoke-test the last "completed" feature. If it's broken, revert and re-open it before touching anything else. Kills the "looks shipped, isn't shipped" bug across sessions.

日本語の概要は準備中です。原文の説明を表示しています。

Archive228/loopkit7552026年7月15日 更新

Archive228 のスキルをすべて見る

このスキルの問題を報告する