Catch the accessibility failures that ship in almost every AI-built UI. Use after building any interactive component.
日本語の概要は準備中です。原文の説明を表示しています。
Negotiate a pre-code contract between generator and evaluator personas that defines what "done" means before any code is written. Turns fuzzy specs into a testable target the evaluator can hold the generator to.
インストールする前に、エージェントに与えられる指示の中身を確認できます。
The evaluator's leverage collapses when "done" is defined after the code exists. The generator ships something, the evaluator finds it plausible, the fuzzy spec silently reshapes to match what got built. Premature victory, dressed up.
Fix the timing: write the contract before the generator writes a line of code. The evaluator's job for the sprint is then mechanical — hold the artifact against the contract, no re-negotiation mid-flight.
Inspired by the planner/generator/evaluator split in Prithvi's March 2026 post on multi-agent harnesses.
shift-notes / a progress file and needs a concrete acceptance target before it picks up the keyboard.claude-progress.txt (or equivalent) under a ## Sprint contract header, with the sprint start timestamp. Both personas reference this exact text for the rest of the sprint.## Sprint contract — <ISO timestamp>
Deliverable: <one sentence, user-observable>
Acceptance predicates:
- [ ] <predicate 1 — script-decidable>
- [ ] <predicate 2>
- [ ] <predicate 3>
Runtime path: <exact URL / CLI / click sequence the evaluator will drive>
Out of scope this sprint:
- <tempting adjacent fix>
- <tempting refactor>
{error: "missing_field"} when name is absent" is.Fresh session, no memory of the previous one. Read shift-notes, find the last signed contract with unchecked predicates. That is your target — no re-planning, no reinterpretation. Drive the runtime path, tick the predicates, ship. If the contract looks wrong on inspection, do not edit it; revert to the planner, get a new one signed.
Two to five minutes of prose before code. In return: the evaluator has something to be strict about, the generator has an anchor against scope drift, and the next session inherits a testable target instead of a vibe.
Trivial edits (typo fix, one-line config change, dependency bump) — the contract overhead exceeds the work. Solo runs with no evaluator persona in the loop — write yourself a one-line acceptance note instead and move on.
まだレビューはありません。使ってみた感想をお寄せください。
概要と使いどころ
Catch the accessibility failures that ship in almost every AI-built UI. Use after building any interactive component.
日本語の概要は準備中です。原文の説明を表示しています。
Before compaction Loopkit extracts decisions into claude-decisions.json (machine-readable). Read it alongside claude-progress.txt at session start — prose is for humans, JSON is for the loop.
日本語の概要は準備中です。原文の説明を表示しています。
Review a diff against the goal spec assuming the code is BROKEN. The reviewer that lives in the maker's head always agrees with itself — this pulls review into a hostile, separate pass. Invoke after every code change before marking work done.
日本語の概要は準備中です。原文の説明を表示しています。
Verify that an endpoint checks ownership, not just authentication. Use on any handler that reads or mutates user data.
日本語の概要は準備中です。原文の説明を表示しています。
Find the exact commit that introduced a bug. Use when something worked before and broke, and you don't know which change did it.
日本語の概要は準備中です。原文の説明を表示しています。
Before picking new work, smoke-test the last "completed" feature. If it's broken, revert and re-open it before touching anything else. Kills the "looks shipped, isn't shipped" bug across sessions.
日本語の概要は準備中です。原文の説明を表示しています。