Use when writing acceptance criteria for a task - express each as an observable Given/When/Then that QA can execute, including negative cases
日本語の概要は準備中です。原文の説明を表示しています。
Test-first discipline and what makes a test worth keeping. Use when implementing any feature or bug fix, before writing production code, and whenever writing or changing a test.
インストール方法を見るインストールする前に、エージェントに与えられる指示の中身を確認できます。
Write the test first. Watch it fail. Write minimal code to pass.
Core principle: If you didn't watch the test fail, you don't know if it tests the right thing.
Violating the letter of the rules is violating the spirit of the rules.
NO PRODUCTION CODE WITHOUT A FAILING TEST FIRST
Wrote code before the test? Delete it. Start over. Don't keep it as "reference", don't "adapt" it while writing tests — delete means delete, implement fresh from the test.
Always: new features, bug fixes, refactoring, behavior changes. Thinking "skip TDD just this once"? Stop. That's rationalization.
Exceptions — no new test required: pure config, copy/text, or asset edits, and generated code you did not hand-write. Say so explicitly in your closing message; this matches the standing acceptance criteria's own exemption, it does not replace it.
expect(build(x)).toBe(build(x)) tests nothing; it recomputes the answer the way the code does, so it passes no matter what. ✅ expect(build({id:1})).toBe("item-1").go test ./pkg/... -run TestName, npx vitest run path/to.test.tsx, flutter test test/x_test.dart, ./mvnw -q test -Dtest=ClassName.Don't assert class names, exact copy, constants, or private structure — a renamed field or a reworded label breaks the test without the behavior changing. Assert the behavior the test's name promises. For pure styling or layout work, the frontend four-width check is the test; a pixel-diff snapshot is not a substitute.
Before you call the step done, ask: if I introduced a wrong constant, flipped a branch, dropped a side effect, returned empty instead of the value, or skipped a validation (zero / empty / nil / unauthorised / malformed) — would ONE of my tests fail? If a mutation like that survives every test green, the test suite has a gap, not full coverage. This is what the run's advisory mutation score is checking for; a high line-coverage number with a low mutation score means assertions are missing, not lines.
Don't delete the existing diff. Write the test the code is missing, then prove it CAN fail: temporarily revert just the covered lines (git stash push <file> or a targeted edit), run the test (MUST FAIL), restore (git stash pop), run again (passes). Then continue from there — this is the same Verify RED discipline applied after the fact, not an exception to it.
Feature: slugify(title) lowercases and hyphenates.
RED test: slugify("Hello World") == "hello-world"
run → FAIL: slugify is not defined ← watched it fail, right reason
GREEN func slugify(s) { return strings.ReplaceAll(strings.ToLower(s), " ", "-") }
run → PASS ← watched it pass
RED test: slugify("A B") == "a-b" (collapse runs)
run → FAIL: got "a--b" ← new behavior, fails first
GREEN collapse whitespace before replacing
run → PASS
REFACTOR extract the whitespace regex, suite still green
Each behavior earned its own failing test first. The second test caught a real gap the first implementation missed — which is the whole point of writing it before the code.
A bug fix starts with a failing test that reproduces the bug. The test proves the fix and prevents regression. Never fix a bug without a reproducing test.
| Excuse | Reality |
|---|---|
| "Too simple to test" | Simple code breaks too. The test takes 30 seconds. |
| "I'll test after" | Tests written after pass immediately and prove nothing. |
| "Already manually tested" | Ad-hoc is not systematic: no record, cannot re-run. |
| "Deleting X hours is wasteful" | Sunk cost fallacy. Unverified code is technical debt. |
| "Keep it as reference" | You will adapt it — that is testing after. Delete it. |
| "TDD will slow me down" | TDD is faster than debugging in review/QA and revision cycles. |
| "Test is hard to write" | Hard to test = hard to use. Simplify the design. |
| "The expectation and the code share the formula, that's fine" | That's a change detector, not a test — it passes no matter what the code does. |
All of these mean: delete the code, start from the test.
Can't check every box? You skipped TDD. Start over.
まだレビューはありません。使ってみた感想をお寄せください。
概要と使いどころ
Use when writing acceptance criteria for a task - express each as an observable Given/When/Then that QA can execute, including negative cases
日本語の概要は準備中です。原文の説明を表示しています。
Use when the diff adds or changes an endpoint, resolver, RPC, job or query that takes an object id, a role check, a request binding or a tenant filter - BOLA/IDOR, function-level authorization, mass assignment and tenant scoping
日本語の概要は準備中です。原文の説明を表示しています。
Use on every UI change - semantic HTML, labels for controls, keyboard-navigable dialogs/menus, visible focus, and never color as the only signal
日本語の概要は準備中です。原文の説明を表示しています。
Use when a task changes any screen, form, dialog, menu or control - Lighthouse/axe scan of the changed screens, a keyboard walk, and the thresholds that fail a task
日本語の概要は準備中です。原文の説明を表示しています。
How to work a task returned with review, QA or UAT findings. Use when a task is in need_revision or PR review comments are in your context.
日本語の概要は準備中です。原文の説明を表示しています。
Use when deciding whether a request needs an analiz task before implementation - the conditions that require the architect's analysis versus going straight to implementation
日本語の概要は準備中です。原文の説明を表示しています。