Catch the accessibility failures that ship in almost every AI-built UI. Use after building any interactive component.
日本語の概要は準備中です。原文の説明を表示しています。
Systematically remove one harness component at a time and measure impact, killing scaffolding that no longer earns its complexity.
インストール方法を見るインストールする前に、エージェントに与えられる指示の中身を確認できます。
Every harness component was added to compensate for a specific model failure. Models improve. Components don't retire themselves. The scaffolding that saved you on Sonnet 4.5 may be dead weight — or actively harmful — on Opus 4.6. Strip it deliberately, one piece at a time, and let evals tell you what still earns its keep.
Inspired by Prithvi's March 2026 harness post on evaluator-generator separation and the general "re-test your assumptions each model bump" discipline.
Inventory the components. List every distinct piece of scaffolding: prompt sections, tool wrappers, post-hoc validators, retry loops, evaluator personas, structured-output enforcers, sandbox rules. One row per component. Note the failure mode each was added to prevent.
Rank by suspicion. Put the components most likely to be obsolete at the top: anything added before the last two model bumps, anything targeting a failure mode you haven't seen recently, anything whose original justification is now folklore.
Pick a baseline eval. You need a repeatable metric before you touch anything. Reuse an existing eval set if you have one; otherwise pick 20–60 tasks representative of production work. Record baseline score, cost, and wall-clock.
Strip one component. Only one. Comment it out or gate it behind a flag — don't delete yet. Re-run the eval.
Compare against baseline.
Commit the delta. Land the strip (or the restore-with-notes) as its own commit. Do not batch multiple strips into one change — you lose the ability to attribute the score movement.
Repeat for the next component. Re-establish baseline from the new state each round, not the original. Compounding strips have compounding effects.
Highest yield in practice:
Don't strip mid-project on a live long-running run — you'll perturb sessions in flight. Do it between projects, or on a forked branch. Also skip if you don't have an eval you trust; stripping without measurement is guessing.
まだレビューはありません。使ってみた感想をお寄せください。
概要と使いどころ
Catch the accessibility failures that ship in almost every AI-built UI. Use after building any interactive component.
日本語の概要は準備中です。原文の説明を表示しています。
Before compaction Loopkit extracts decisions into claude-decisions.json (machine-readable). Read it alongside claude-progress.txt at session start — prose is for humans, JSON is for the loop.
日本語の概要は準備中です。原文の説明を表示しています。
Review a diff against the goal spec assuming the code is BROKEN. The reviewer that lives in the maker's head always agrees with itself — this pulls review into a hostile, separate pass. Invoke after every code change before marking work done.
日本語の概要は準備中です。原文の説明を表示しています。
Verify that an endpoint checks ownership, not just authentication. Use on any handler that reads or mutates user data.
日本語の概要は準備中です。原文の説明を表示しています。
Find the exact commit that introduced a bug. Use when something worked before and broke, and you don't know which change did it.
日本語の概要は準備中です。原文の説明を表示しています。
Before picking new work, smoke-test the last "completed" feature. If it's broken, revert and re-open it before touching anything else. Kills the "looks shipped, isn't shipped" bug across sessions.
日本語の概要は準備中です。原文の説明を表示しています。