accounts
無料Manage multiple Claude Code accounts: add, list, check, launch, and install shell aliases for 10+ isolated CLAUDE_CONFIG_DIR profiles.
日本語の概要は準備中です。原文の説明を表示しています。
Testing: TDD, E2E, preferred patterns, test-value audits, verification, agent testing.
インストール方法を見るインストールする前に、エージェントに与えられる指示の中身を確認できます。
Seven modes. Match the request to one mode and follow its section. Read repository CLAUDE.md first -- project conventions override defaults here.
| Request matches | Go to |
|---|---|
| Write tests first, TDD, red-green-refactor | TDD |
| Flaky, brittle, test smell, over-mocking, slow tests | Pattern Quality |
| Audit or prune low-value, duplicate, or implementation-coupled tests | Test Audit |
| Test an agent, subagent testing, validate agent | Agent Testing |
| Run vitest, JavaScript/TypeScript tests | Vitest Runner |
| Playwright, E2E, end-to-end, browser test | E2E (Playwright) |
| Verify completion, final check, defense in depth | Verification |
RED-GREEN-REFACTOR cycle with strict phase gates. Each feature gets its own cycle. Do not batch multiple features into one cycle.
Write a test describing desired behavior before implementation exists. Use Arrange-Act-Assert, descriptive names, one concept per test. Run the test and show full output.
Every new or changed test, in any mode, must first pass the authoring gate in
references/audit-test-value.md.
Gate -- proceed only when all true:
If test passes before implementation: assertions are too weak, or the feature already exists. If test fails for wrong reason (syntax, import, setup): fix those first, then re-run until it fails for the right reason.
Write ONLY enough code to make the failing test pass. No extra features. Hardcoded values are acceptable initially. Run the test and the full suite; show complete output.
Gate -- proceed only when all true:
Improve code quality without changing behavior. Establish a green baseline, refactor incrementally, run tests after every step. Test behavior, not internals.
Gate -- proceed only when all true:
Commit test and implementation as an atomic unit. Run the full suite first.
| Symptom | Cause | Fix |
|---|---|---|
| Test passes in RED phase | Weak assertions or feature exists | Strengthen assertions; check for existing implementation |
| Wrong failure reason | Setup incomplete, missing deps | Fix syntax/imports first, re-run |
| Tests green but feature broken | Tests miss actual usage | Add integration tests; test with real data |
| Refactoring breaks tests | Tests coupled to internals | Test behavior not implementation; refactor in smaller steps |
Identify and fix testing mistakes across unit, integration, and E2E suites. Test behavior, be reliable, run fast, fail for the right reasons.
Locate test files (*_test.go, test_*.py, *.test.ts, *.spec.js). Scan
for these 10 failure modes:
| # | Pattern | Detection Signal |
|---|---|---|
| 1 | Testing implementation details | Asserts on private fields, spy on private methods |
| 2 | Over-mocking / brittle selectors | Mock setup > 50% of test code, CSS nth-child |
| 3 | Order-dependent tests | Shared mutable state, numbered test names |
| 4 | Incomplete assertions | != nil, > 0, toBeTruthy(), no value checks |
| 5 | Over-specification | Exact timestamps, hardcoded IDs, asserting defaults |
| 6 | Ignored failures | @skip, .skip, xit, empty catch, _ = err |
| 7 | Poor naming | testFunc2, it('works'), it('handles case') |
| 8 | Missing edge cases | Only happy path, no empty/null/boundary/error tests |
| 9 | Slow test suites | Full DB reset per test, no parallelization |
| 10 | Flaky tests | sleep(), time.Sleep(), unsynchronized goroutines |
Document each finding with file:line, severity, issue, and impact.
Gate: At least one quality issue identified with file:line reference.
Fix one pattern at a time. Preserve test intent. Prevent over-engineering.
Gate: Findings ranked. User agrees on fix scope.
For each issue (highest priority first): show current code, show fixed code, apply fix, run tests. Guide toward behavior testing:
_getUser() -> test what happens when a user exists or notRun the specific fixed test first, then the full file or package. If a fix breaks a previously-passing test, investigate before proceeding.
Gate: Each fix verified. Tests pass after each change.
Run full suite. Verify flaky tests are now deterministic (run 3x). Confirm no no tests were accidentally removed or disabled. Report: bad patterns fixed, files modified, tests affected, suite status.
Gate: Full suite passes. Summary delivered.
| Problem | Fix |
|---|---|
| Cannot determine if pattern is a quality issue | Check comments, consider test layer, flag MEDIUM with trade-offs |
| Fix changes test behavior | Identify original intent, write correct assertion, note as separate finding |
| Suite has hundreds of quality issues | Fix HIGH severity first, recommend TDD going forward, suggest fix-on-touch |
Find and remove tests that do not earn their maintenance cost, and the
test-only production seams they keep alive. Load
references/audit-test-value.md and follow it; it holds the authoring gate,
junk patterns, retention bar, evidence fields, validation, and handoff.
Use Pattern Quality to fix a test that guards real behavior badly. Use Test Audit to decide whether a test should exist at all.
| Scope | Sub-mode |
|---|---|
| Writing or changing a test | Authoring gate |
| Focused sweep of a few candidates | Audit |
| Every test one subsystem owns | Campaign |
Gate: Every deletion has all candidate evidence fields recorded. Owner and full suites pass. Production and test LOC reported separately.
TDD methodology applied to agent development. Test what the agent DOES, not what the prompt SAYS. Each test runs in a fresh subagent to avoid context pollution.
| Agent Type | Min Tests | Coverage |
|---|---|---|
| Reviewer | 6 | 2 real issues, 2 clean, 1 edge, 1 ambiguous |
| Implementation | 5 | 2 typical, 1 complex, 1 minimal, 1 error |
| Analysis | 4 | 2 standard, 1 edge, 1 malformed |
| Routing/orchestration | 4 | 2 correct route, 1 ambiguous, 1 invalid |
No agent is simple enough to skip testing.
Read the agent file and referenced skills. Extract testable claims (inputs, output structure, routing triggers, error conditions). Write a test plan to a file. Dispatch subagent via Task tool with test inputs. Capture results verbatim. Identify failure patterns.
Gate: All cases executed. Outputs captured. Failures documented.
Prioritize failures by severity. Make one fix at a time. Re-run ALL test cases after each fix. If a fix causes regression, revert and try a different approach.
Gate: All cases pass. No regressions.
Add edge case tests (empty, large, unusual, ambiguous inputs). Run consistency tests (same input 3x; outputs should have same structure and key findings). Run full regression suite.
Gate: Edge cases handled. Consistency verified. Full suite green.
Run existing Vitest tests and report results. A check-only request does not authorize changing tests, assertions, dependencies, or configuration.
Check package.json, vitest.config.*, and vite.config.* to confirm Vitest.
Use the installed project version; avoid implicit npx downloads. If Vitest is
unavailable, report setup needed rather than installing it. Always use run;
bare vitest enters watch mode.
| Scope | Command |
|---|---|
| Full suite | npx vitest run --reporter=verbose 2>&1 |
| File or directory | npx vitest run path/to/test.ts 2>&1 |
| Test-name pattern | npx vitest run -t "pattern" 2>&1 |
| Coverage | npx vitest run --coverage 2>&1 |
Capture exit code and full output. Report pass/fail, scope, counts, duration. For failures: retain file, test name, assertion diff, relevant stack. Nonzero exit is failure; partial output is not a passing run.
| Problem | Fix |
|---|---|
| Vitest missing / no node_modules | npm install or npm install -D vitest |
| No test files found | Check naming (*.test.ts, *.spec.ts) and include/exclude globs |
| Missing DOM environment | Check for jsdom/happy-dom in config; suggest devDependency |
| Out of memory | Batch by directory, use --pool=forks or --shard=1/N |
| Failing assertions | Report mismatch; if fixing authorized, determine whether implementation or test is wrong |
Playwright-based E2E testing: Scaffold, Build, Run, Validate. Each phase produces an artifact and must pass its gate.
@playwright/test installed: npx playwright --version. If missing:
npm install -D @playwright/test && npx playwright install.tests/e2e/{auth,features,api}/, pages/,
artifacts/{screenshots,traces,videos}/.playwright.config.ts. Bake in failure diagnostics: screenshot: 'only-on-failure',
trace: 'on-first-retry', video: 'retain-on-failure'. CI retries:
retries: process.env.CI ? 2 : 0.npx tsc --noEmit.Gate: playwright.config.ts exists AND tests/e2e/ exists.
Write POM classes in pages/ for each feature area. All locators use
data-testid via page.getByTestId(). No inline locators in spec files.
Write spec files in tests/e2e/<area>/. Verify: npx tsc --noEmit.
Gate: At least one .spec.ts under tests/e2e/ AND npx tsc --noEmit
exits 0.
BASE_URL).npx playwright test.--repeat-each=5 to distinguish flaky from broken.test.fixme() and a tracking TODO.
Never delete a failing test. Use test.skip() only for environment guards.Gate: playwright-results.json exists and parses as valid JSON.
unexpected
and flaky entries.e2e-report.md.Gate: e2e-report.md exists.
| Symptom | Fix |
|---|---|
npx tsc --noEmit fails | Check @playwright/test in devDeps, verify tsconfig includes test dir |
| Pass locally, fail CI | npx playwright install --with-deps in CI; verify BASE_URL |
| Results JSON missing | Check JSON reporter in config; check for OOM/process kill |
| Locator timeout on existing element | await expect(locator).toBeVisible() before interaction; check overlays |
fill() appends | locator.clear() then locator.fill() |
| Flaky (4/5 pass) | Quarantine with test.fixme(), reproduce with --repeat-each=10, check missing waitFor |
Confirm flaky vs. broken: --repeat-each=5 --retries=0. If fails at least once
in 5, it is flaky. Fix if root cause is clear; quarantine otherwise. Verify fix
with --repeat-each=10 --retries=0 (must pass 10/10).
Defense-in-depth verification before declaring any task complete. Match checks to affected behavior and repository requirements.
git status --short and git diff. Read changed code;
check imports, error handling, compatibility, unintended edits.pass is not automatically a stub.| Language | Tests | Build/syntax | Lint |
|---|---|---|---|
| Python | pytest -v | python -m py_compile {files} | ruff check {files} |
| Go | go test ./... -v -race | go build ./... | golangci-lint run ./... |
| JavaScript | npm test | npm run build | npm run lint |
| TypeScript | npm test | npx tsc --noEmit | npm run lint |
| Rust | cargo test | cargo build | cargo clippy |
Reuse a passing result when it covers the current task, checked files, dependencies, and environment. Keep its command, scope, state, and log path. After edits, rerun affected checks. Do not claim inherited results as your own. Required CI checks still apply to the delivered commit.
| Problem | Fix |
|---|---|
| No tests | Manual checks; state coverage gap; add regression test if warranted |
| Missing dependencies | Use repo environment; report missing tool; unrun checks are not passes |
| Build/test failure | Retain failing command and diagnostic; identify cause, fix, rerun |
| Missing wiring or data flow | Name where integration stops; repair it |
| Rationalization | Required Action |
|---|---|
| "I loaded the patterns, that's enough" | Loading is not applying. Check against patterns at each gate. |
| "This task is simple, full rigor is overkill" | Apply proportionate rigor, never zero. |
| "The gate basically passes" | Either it passes with evidence or it does not. |
Completion self-check: Did I verify or assume? Did I run tests or just read code? Did I complete everything or just the "important" parts? Can I show evidence?
Load when the signal applies.
| Signal | Load | Content |
|---|---|---|
| TDD phase steps, language commands | references/tdd-phase-guidance.md | RED-GREEN-REFACTOR steps per language |
| TDD walkthroughs | references/tdd-examples.md | Go, Python, JavaScript worked examples |
| BAD/GOOD code per failure mode | references/patterns-preferred-pattern-catalog.md | Code examples per pattern per language |
| Failure mode classification | references/patterns-quality-catalog.md | 10 failure mode descriptions |
| Language-specific fix strategies | references/patterns-fix-strategies.md | Fix patterns and tooling per language |
| Test blind spots | references/patterns-blind-spot-taxonomy.md | 6-category gap taxonomy |
| Load test scenarios | references/patterns-load-test-scenarios.md | Smoke, stress, spike, soak configs |
| New-test gate, test audits, pruning campaigns | references/audit-test-value.md | Authoring gate, junk patterns, retention bar, evidence, validation |
| Agent dispatch patterns | references/agents-testing-patterns.md | Dispatch, negative, A/B |
| Agent testing examples | references/agents-examples-and-errors.md | Worked examples and error cases |
| E2E async patterns | references/e2e-async.md | Promise.all, race conditions, teardown |
| E2E auth testing | references/e2e-auth.md | Login, storageState, OAuth, SSO, JWT |
| E2E config templates | references/e2e-templates.md | playwright.config.ts, POM, CI/CD |
| E2E POM and waiting | references/e2e-playwright-patterns.md | POM examples, multi-browser |
| E2E Web3 wallet | references/e2e-wallet-testing.md | MetaMask testing patterns |
| E2E financial flows | references/e2e-financial-flows.md | Payment flow testing |
| Stub detection | references/verify-adversarial-methodology.md | Four-level checks, goal-backward verification |
| Domain checklists | references/verify-checklist.md | Schema change, compatibility checks |
| Verification examples | references/verify-verification-examples.md | Bug fix, refactor, migration walkthroughs |
@skip, @ignore, xit, .skip without expiration datetime.sleep(), setTimeout() in test codetest1, test2)!= nil, > 0, toBeTruthy() without value checksStrict TDD prevents most quality issues: RED catches incomplete assertions, GREEN minimum prevents over-specification, watching failure confirms you test behavior not mocks, incremental cycles prevent interdependence, refactor phase reveals implementation coupling.
まだレビューはありません。使ってみた感想をお寄せください。
概要と使いどころ
Manage multiple Claude Code accounts: add, list, check, launch, and install shell aliases for 10+ isolated CLAUDE_CONFIG_DIR profiles.
日本語の概要は準備中です。原文の説明を表示しています。
Improve architecture across modules by deepening interfaces.
日本語の概要は準備中です。原文の説明を表示しています。
Assessment: read-only inspection, codebase overview, value analysis, health checks, ADR consultation, decision analysis, multi-perspective critique.
日本語の概要は準備中です。原文の説明を表示しています。
Background memory consolidation — overnight review, merge, and injection payload for memory files.
日本語の概要は準備中です。原文の説明を表示しています。
Jev-driven browser automation: Jev picks operations, programs execute, a text model writes field values only when Jev cannot pick one from the goal.
日本語の概要は準備中です。原文の説明を表示しています。
Write, compose, integrate, and improve programs that call Jev, TypeSafe's System One judgment model.
日本語の概要は準備中です。原文の説明を表示しています。