Catch the accessibility failures that ship in almost every AI-built UI. Use after building any interactive component.
日本語の概要は準備中です。原文の説明を表示しています。
Build a repeatable eval loop that grades agent output with an LLM judge, so prompt/skill changes get scored against a baseline instead of eyeballed. Reuses loopkit's verifier subagent as the grader — do not build a new one.
インストールする前に、エージェントに与えられる指示の中身を確認できます。
Every prompt tweak in a long-running agent looks like an improvement in the moment. The only way to know is a graded run against fixed inputs. Loopkit already ships .claude/agents/verifier.md — that is your grader. Do not rebuild it.
inputs.jsonl → runner → outputs.jsonl → verifier (per row) → verdicts.jsonl → diff vs baseline
Each stage writes to disk. No stage holds the whole run in context.
One JSON object per row: {"id": "case-01", "input": "...", "expected": "..."}.
inputs-v2.jsonl) when you change it. Never edit in place — you lose the baseline.A dumb loop: for each row, call the model with the current prompt/skill, capture output, write {"id": ..., "output": ...} to outputs.jsonl. No grading here — just capture.
If the runner is smart it will bias the eval. Keep it dumb.
Fan out one subagent per row (see subagent-fanout). Each gets:
.claude/agents/verifier.md.Verifier returns strict JSON: {"pass": bool, "why": "..."}. Collect into verdicts.jsonl.
Two runs of the same eval on two prompt versions → compare pass rates per case. What matters:
A change that raises the mean but adds regressions is usually a loss — the new failures are cases you already knew worked.
verdicts.jsonl to git.The verifier is already yours. The harness is 100 lines of glue around it.
まだレビューはありません。使ってみた感想をお寄せください。
概要と使いどころ
Catch the accessibility failures that ship in almost every AI-built UI. Use after building any interactive component.
日本語の概要は準備中です。原文の説明を表示しています。
Before compaction Loopkit extracts decisions into claude-decisions.json (machine-readable). Read it alongside claude-progress.txt at session start — prose is for humans, JSON is for the loop.
日本語の概要は準備中です。原文の説明を表示しています。
Review a diff against the goal spec assuming the code is BROKEN. The reviewer that lives in the maker's head always agrees with itself — this pulls review into a hostile, separate pass. Invoke after every code change before marking work done.
日本語の概要は準備中です。原文の説明を表示しています。
Verify that an endpoint checks ownership, not just authentication. Use on any handler that reads or mutates user data.
日本語の概要は準備中です。原文の説明を表示しています。
Find the exact commit that introduced a bug. Use when something worked before and broke, and you don't know which change did it.
日本語の概要は準備中です。原文の説明を表示しています。
Before picking new work, smoke-test the last "completed" feature. If it's broken, revert and re-open it before touching anything else. Kills the "looks shipped, isn't shipped" bug across sessions.
日本語の概要は準備中です。原文の説明を表示しています。