Audit or rewrite AGENTS.md so it holds only lasting principles. Run it only when the user asks for it by name; never invoke it on your own.
日本語の概要は準備中です。原文の説明を表示しています。
Write tests that pin real behavior instead of implementation details, config values, or lucky samples. Use when adding tests for new behavior, writing a regression test, fixing a brittle or flaky test, reviewing a test diff, or when a test breaks after a refactor or config change that didn't change behavior.
インストール方法を見るインストールする前に、エージェントに与えられる指示の中身を確認できます。
A good test fails only when real behavior breaks, and passes through every refactor or config change that preserves it. Most bad tests fail the opposite way: red on harmless changes, green while the real path is broken. Every rule below serves that one goal.
For deciding whether coverage adds independent proof, where it belongs, or which existing tests can go, use audit-tests.
length == 1 passes even when normalization is broken; also assert
the stored value equals the expected canonical form.skipIf/conditionals — a skipped
assertion hides the coupling and stops covering the path. Feed the function
fixed inputs, or stub the lookup so fixture ids resolve to fixed values.
Litmus: "would this break if a config value changed with no logic change?"A suite that grabs the desktop is a suite people stop running. Whatever the test needs, it takes the least intrusive form of it:
makeKeyAndOrderFront, no
raising, no moving the pointer, no simulated clicks into the session, no audio
playback. A UI a person must not be interrupted by is still fully testable:
build the window, drive its model, and assert on the rendered view. Where a
screenshot is the evidence, capture the window's own image offscreen rather
than photographing the screen. Keep activation on the code path a person
triggers, and exercise it by calling the handler, not by launching the app
repeatedly.A test you never saw fail is decoration. For any regression test — especially one written after the fix — falsify it once: revert or break the production code the way the bug would, confirm red for the expected reason, restore, confirm green. Verify the revert actually took: a stash or checkout with a wrong pathspec reverts nothing, silently, and the "red" run quietly tests the fixed code. The tell: the "red" numbers equal the green numbers.
Triage in order of likelihood before editing anything:
Never tune constants to make one test pass without rerunning the neighbors: coupled systems reshuffle. If two consecutive tweaks each break different tests, stop poking — the control surface is wrong; find the mechanism.
Walk this on any test diff, apply fixes in the same pass, re-run the suite:
skipIf-conditional on config? → mock the seam.まだレビューはありません。使ってみた感想をお寄せください。
概要と使いどころ
Audit or rewrite AGENTS.md so it holds only lasting principles. Run it only when the user asks for it by name; never invoke it on your own.
日本語の概要は準備中です。原文の説明を表示しています。
Audit the choices an implementing agent made, not its diff — a pure decision audit that traces the session's history into a choices ledger, changes no code, and never blocks an unsupervised run. Working code still embeds architecture the user never chose; surface it because future work inherits it. Use when the user wants to review the decisions the AI made on their behalf, before merging or committing AI-implemented work, when integrating a delegated subagent's pass, or when a fix "works" but might be a point fix.
日本語の概要は準備中です。原文の説明を表示しています。
Audit and prioritize performance work through bounded-work and forward-progress checks. Use when asked to find performance issues, investigate freezes/500s/OOMs, review retries or polling for no-progress loops, rank a performance backlog, or implement the next simple performance fixes.
日本語の概要は準備中です。原文の説明を表示しています。
Audit whether tests earn their maintenance cost and which suite owns each contract. Use when pruning redundant tests, investigating implementation coupling or test-only production hooks, reviewing the value of proposed coverage, or auditing an entire subsystem. Use write-tests to implement the resulting test changes.
日本語の概要は準備中です。原文の説明を表示しています。
Run fast, progressive experiments to map tunable parameters and their effects. Use when optimizing code, prompts, configurations, or other artifacts against a goal or benchmark, or when a long research run needs a planned hypothesis queue, focused trials, and durable findings.
日本語の概要は準備中です。原文の説明を表示しています。
Use Claude Code as an independent `claude -p` subagent when the user explicitly asks for Claude, wants a second-agent opinion from Claude, or asks to delegate a well-scoped task to Claude. Supports selecting `--model` and thinking/effort level with defaults of `opus` and `high`.
日本語の概要は準備中です。原文の説明を表示しています。