Use when reviewing UI for accessibility — WCAG 2.2 AA, keyboard nav, focus, ARIA, contrast, screen-reader semantics — even on 'is this a11y-OK?' or 'mach das barrierefrei'.
日本語の概要は準備中です。原文の説明を表示しています。
Implementing a feature, fixing a bug, refactoring — failing test first, then the code. For a WRONG test, `testing-anti-patterns` wins.
インストール方法を見るインストールする前に、エージェントに与えられる指示の中身を確認できます。
Do NOT use when:
.md, AGENTS.md, README)agent-config package on skill/rule markdownPlain TDD (red → green → refactor) is the right size for most work. A small subset benefits from a gated 5-phase wrapper — Spec → Pseudocode → Architecture → Refine → Complete — where each gate produces a written artifact before the next phase runs.
Decision tree — escalate if any branch is true:
When escalating: drop the Spec artifact in agents/roadmaps/ (or the
project's planning location), capture it as an ADR via adr-create
when the decision is load-bearing, then run plain TDD inside each
Refine cycle. Do not skip RED→GREEN inside a SPARC phase — the
wrapper adds gates, not exemptions.
For everything else (single-AC ticket, leaf-module change, bug fix), stay on plain TDD — the section above.
1. Write ONE failing test that describes the desired behavior.
2. Run it. WATCH it fail for the right reason.
3. Write the MINIMUM production code to make it pass.
4. Run it again. Watch it pass.
5. Clean up (rename, deduplicate) while keeping the test green.
If step 2 is skipped, the test is not trusted — a test that has never failed proves nothing about the code under test.
UNTESTED CODE THIS TASK JUST WROTE, AND A TEST IS NEEDED — DELETE THE CODE,
WRITE THE TEST, REIMPLEMENT. NEVER KEEP IT "AS REFERENCE".
THIS LAW COVERS YOUR OWN UNTESTED OUTPUT. IT IS NOT A LICENCE TO DELETE
PRE-EXISTING CODE, AND NEVER OVERRIDES A REUSE VERDICT.
Reading the existing implementation while writing its test is test-after-the-fact with extra steps. Which code that applies to has three answers, and only the first is a deletion:
| The code is | Do |
|---|---|
| untested, written by this task | delete it, write the test, reimplement — the Iron Law above |
| pre-existing and tested | keep it. Its tests are the record of its behaviour; deleting it to re-derive the same thing discards evidence and contradicts the reuse verdict |
| pre-existing and untested | do NOT delete. Write a characterization test pinning the behaviour it has today — including the behaviour you think is wrong — then change it under that test |
The middle row is the one this law used to get wrong: unqualified, it read as
a standing instruction to delete tested legacy that a reuse verdict would
keep. Case detail, the characterization-test procedure, and the 12-row
anti-rationalization table are in
testing-anti-patterns/process-anti-patterns.md,
which keeps this skill under the 400-line sunset trigger.
The flow runs as four modes. Each Forbidden item names how a reviewer checks it from the diff — an unverifiable prohibition does not ship.
| Mode | Goal | Activities | Forbidden (diff check) | Output contract |
|---|---|---|---|---|
| Design | One-sentence behavior + enumerated cases | Steps 1–2 | No production code (diff touches no src/** production path) · no test bodies yet | Case list (happy/boundary/error) |
| Test-Red | A failing test that fails RIGHT | Steps 3–4 | No production edits (diff = tests/** only) · the failure must be about the behaviour under test. Valid: a failing assertion · a missing target — class-not-found, or a compile/type error naming the unimplemented symbol · a contract failure (wrong shape, wrong status, unmet interface). Invalid: a broken fixture · a syntax error in the test · a missing unrelated dependency · a runner or environment fault | Failing test + its observed failure, named as one of the three valid classes |
| Implement | Minimum code to green | Steps 5–6 | No test edits (no tests/** paths in Implement-phase diffs — changing the assertion to fit the code is the canonical violation; genuinely-wrong test → STOP and ask, never silently edit) · no scope beyond the one case | Green run output |
| Debug | Fix a defect found later | (re-enter at 3) | No bugfix before a reproducing regression test exists (the fix commit contains a tests/** addition that fails without the fix) | Regression test + fix, verified red→green |
The discriminator is whether the failure is about the behaviour under test, never where in the run it surfaces. A class that does not exist yet can only fail at load, so demanding an assertion would force a production stub before the first test — the exact thing this skill forbids. The four invalid classes are failures of the harness: they would fail identically with the behaviour fully implemented, so they measure nothing about it. Unsure → re-read the failure output and name which of the seven it is; an unclassified red is not a RED.
Step 0 of any resume: infer the mode from observable state — never assume Design:
| Observed state | Resume in |
|---|---|
| No test for the target behavior | Design |
| Test exists, currently failing at an assertion | Implement |
| Test exists + passing, defect reported | Debug |
| Test failing at load because the target does not exist yet | Implement (that is a valid RED) |
| Test failing on a harness fault — fixture, syntax, unrelated dependency, runner | Test-Red (fix the test, not the code) |
At every mode transition, one consent-checkpoint sentence (per
ask-when-uncertain / autonomous-execution — no new mechanism): name the
mode you are leaving, the output contract you hand over, and the mode you
enter; under an autonomous mandate the sentence is stated, not asked.
State in one sentence: "When X happens, the system should do Y."
If you cannot state it in one sentence, the scope is too big — split into multiple tests, each covering one sentence.
Before the first test is written, run the
test-case-discovery funnel for the
behavior: dimension scan → case synthesis → optional subagent cross-check →
prioritization. Do not proceed to step 3 with fewer than the floor:
Each case from the list then gets its own RED → GREEN cycle (steps 3–6). A behavior whose only test is the happy path is not done — it is the first item of an unfinished case list.
Write the smallest test that expresses the sentence from step 1.
it_rejects_empty_email, not test_email_1.Execute the single test (targeted, not the full suite):
# PHP/Pest
./vendor/bin/pest --filter=it_rejects_empty_email
# JS/Vitest
npx vitest run --testNamePattern "rejects empty email"
Required observations before proceeding:
Add just enough production code to make the test green. No extra features, no unrelated refactoring, no "while I'm here" cleanups.
If you feel the urge to add a parameter, edge case, or helper not covered by the current test — stop. That belongs in the next RED step, not this GREEN step.
Re-run the same targeted command. Required:
With all tests green, you may:
Do not add new behavior during refactor — that needs its own failing test first. Re-run tests after the refactor to confirm still-green.
Back to step 1 with the next single-sentence behavior.
Twelve common rationalizations that fire before the test is written —
plus the delete-and-restart Iron Law — live in
testing-anti-patterns/process-anti-patterns.md.
Read the table when:
For mock-isolation failure modes (separate concern), see
testing-anti-patterns.
// tests/Unit/EmailValidatorTest.php — RED
it('rejects empty email', function () {
$result = (new EmailValidator())->validate('');
expect($result->isValid())->toBeFalse();
expect($result->error())->toBe('Email required');
});
Run: ./vendor/bin/pest --filter='rejects empty email' → fails
(EmailValidator does not exist yet, or returns isValid()=true).
// app/Validators/EmailValidator.php — GREEN (minimum)
final class EmailValidator
{
public function validate(string $email): EmailResult
{
if (trim($email) === '') {
return EmailResult::invalid('Email required');
}
return EmailResult::valid();
}
}
Run the filter again → passes. No additional rules (format, MX, length) until a next failing test drives them.
// src/retry.test.ts — RED
import { retry } from './retry';
it('retries a failing operation up to 3 times', async () => {
let attempts = 0;
const op = async () => {
attempts += 1;
if (attempts < 3) throw new Error('transient');
return 'ok';
};
await expect(retry(op)).resolves.toBe('ok');
expect(attempts).toBe(3);
});
Run: npx vitest run --testNamePattern "retries a failing" → fails
(retry is undefined).
// src/retry.ts — GREEN (minimum)
export async function retry<T>(op: () => Promise<T>): Promise<T> {
let lastError: unknown;
for (let i = 0; i < 3; i += 1) {
try { return await op(); } catch (e) { lastError = e; }
}
throw lastError;
}
Run again → passes. Configurable attempt count, backoff, and jitter all wait for their own failing tests.
The discipline above is stack-independent; the command is not, and every
runner invocation elsewhere in this file is a PHP one. Resolve the runner with
resolve_toolchain (work_engine/stack/runner.ts) and read its ecosystems,
then filter with that ecosystem's own form — never the whole suite:
ecosystems | Filtered probe |
|---|---|
php | vendor/bin/pest --filter '<name>' · php artisan test --filter '<name>' |
js | npx vitest run -t '<name>' · npx jest -t '<name>' |
python | pytest -k '<expr>' · pytest path/to/test.py::test_name |
go | go test -run '<regexp>' ./pkg/... |
A TypeScript repository resolves the ecosystem js, never typescript. The
whole suite is the final gate and never the per-iteration probe, in every row.
expect() with three or four assertions on unrelated fields describes
multiple behaviors. Split them.it('works') — no behavior describedtest-case-discoveryquality-toolspest-testing/tests:executesystematic-debuggingBefore marking TDD work complete:
See also developer-like-execution
for the broader think → analyze → verify loop this skill plugs into.
まだレビューはありません。使ってみた感想をお寄せください。
概要と使いどころ
Use when reviewing UI for accessibility — WCAG 2.2 AA, keyboard nav, focus, ARIA, contrast, screen-reader semantics — even on 'is this a11y-OK?' or 'mach das barrierefrei'.
日本語の概要は準備中です。原文の説明を表示しています。
Use when defining or auditing the activation event — aha-moment selection, retention correlation, falsifiable definition. Triggers on 'what is our aha moment', 'redefine activation'.
日本語の概要は準備中です。原文の説明を表示しています。
Use when capturing an architectural decision — file naming, next ADR number, Status / Context / Decision / Consequences, index regen; fires even without saying 'ADR'.
日本語の概要は準備中です。原文の説明を表示しています。
Adversarial critique — devil's advocate, stress-test, honest teardown ('poke holes', 'be brutal', 'was hältst du davon'); explicit request only. Routine code or design review → code-review.
日本語の概要は準備中です。原文の説明を表示しています。
Use when reading, creating, or updating agent documentation, module docs, roadmaps, or AGENTS.md. Understands the full .augment/, agents/, and copilot-instructions structure.
日本語の概要は準備中です。原文の説明を表示しています。
Use for an adversarial red-team / blue-team / auditor review of an AI agent's CONFIG + behaviour (rules, skills, MCP, hooks, permissions) — attack-chain → defensive-gap list, not a code audit.
日本語の概要は準備中です。原文の説明を表示しています。