Audit or rewrite AGENTS.md so it holds only lasting principles. Run it only when the user asks for it by name; never invoke it on your own.
日本語の概要は準備中です。原文の説明を表示しています。
Implement an existing spec through committed passes. Use for long or multi-pass specs that need maintenance checkpoints to periodically clean code, handoffs, priorities, and plan bloat before drift accumulates.
インストール方法を見るインストールする前に、エージェントに与えられる指示の中身を確認できます。
Build the active spec to completion, one reviewable pass at a time. The spec is the implementation plan; user constraints and accepted experimental evidence remain binding when the architecture changes.
A pass (usually one slice) is a commit checkpoint, not a stopping point. The job is the whole spec — every slice, every global TODO — not the first green commit. Finishing a pass means starting the next one, not handing back to the user. Only stop when the spec is fully implemented (or a genuine blocker needs a decision only the user can make).
Work in parallel wherever the graph allows. Do not walk the ladder one slice at a time when slices are independent. Read the spec's dependency graph as a wavefront and delegate independent passes to subagents that run concurrently (see Rules) — you orchestrate and integrate; only serialize what genuinely depends on prior work.
Read the repo README, the spec README, and the next slice before editing. Load any skills named by the spec. Identify the current pickup point, global TODOs, required gates, and what must stay green. If a multi-slice spec lacks a live handoff prompt, add one before the first pass ends. For a validated spike, apply Preserve the winner below before editing.
Reconcile the plan with the current code. If the slice would preserve a development-only shim, duplicated type, weak wrapper, or obsolete path, replace it with the simpler architecture and update the spec handoff.
Implement one coherent pass: usually one slice, one vertical checkpoint, or one architecture correction. Keep the review surface small enough to audit.
Verify the actual contract. For behavior changes, run the focused unit tests first; for browser-visible work, use the real browser/harness and inspect screenshots so the subject is framed and readable, not merely nonblank. A visual CHANGE carries two extra proofs before you claim it: a byte/pixel diff against the pre-change baseline on the production route (static checks and re-anchored assertions all pass on a no-op — a palette pass once shipped "verified" while the production frame was byte-identical), and an unprimed screenshot-critique — at the user's reported framing when the pass answers their visual bug report — before declaring it fixed; the implementer's eyes are primed by the fix and repeatedly pass what fresh eyes catch. Never weaken an existing default gate or repin a failing contract without proving the old contract is wrong.
Review the change list and clean up after every pass, before committing.
Read git status/git diff --stat line by line and account for every path:
one-off probes, shot scripts, scratch files, nohup.out, ad-hoc screenshot
dirs, and disposable SPIKE/debug notes never enter a commit — scratch stays
out of the tree. Preserve frozen references, fixtures, and research evidence
in the project's reference or spec assets area. A file without a durable
purpose does not ship.
Delegated agents leak these; the integrating reviewer re-checks the merged
tree with the same eye.
Sizing the pass — measuring what it cost and reworking it when the line count is out of proportion to the behaviour it delivers — belongs to the shape pass in step 6, under refactor-clean.
Run review at the end of every pass, before committing.
It sequences the three closeout lenses — refactor-clean on the shape,
code-review on the settled diff, write-docs on what the change touched — and
you apply the fixes from all three. Long specs are where sediment compounds,
so hold the shape pass to this pass's own output: dev-only shims, duplicated
concepts, parallel abstractions, and compatibility wrappers it introduced
collapse into the clean contract with one owner, so the code reads as designed
today, not tacked on. Last, run
audit-choices on the cleaned pass — a pure
audit that appends every decision made where the spec was silent (your own,
and each delegated subagent's when integrating) to the spec's choices
ledger (specs/<feature>/choices.md) with verdicts, changing no code
itself. You act on its findings: redo unsound choices from their corrected
decisions, adopt the recorded provisional call on any user-only entry — the
audit never blocks the run. Rerun the affected checks, then commit only the
focused changes from this pass.
Update the spec README's "Next Agent Prompt": status, completed work, next pickup point, blockers, changed gates, and any architecture decision that changed the plan.
Run a maintenance checkpoint as part of the loop, not as endgame cleanup. Trigger it after a red pass, after every two or three slice commits, after a rebase/resume/compaction, before changing feature areas, when evidence invalidates the plan, when the handoff contradicts the TODO/graph, when the choices ledger's entries cluster around one slice, or when the active prompt grows hard to scan. Long specs bloat repeatedly; cleanup is a normal pass, not a cosmetic chore.
A checkpoint cleans both plan and code before more feature work:
Commit the checkpoint as its own focused pass when cleanup changes the spec, code shape, or handoff enough that future agents would otherwise inherit stale context. It is done only when a fresh agent can read the README handoff, TODO, and slice graph and choose the same next action without conversation history.
Continue. If any slice or global TODO is still open, go straight back to step 1 for the next one — same session, no pause for acknowledgement. Keep looping until every TODO is closed.
Run review once more over the whole spec. The per-pass reviews each judged one slice against the code as it stood then; this one judges the finished feature. Scope it to the spec's full diff, not the last pass — that is the only scope where duplication spread across slices, a shape that only reads wrong once every slice has landed, and docs that describe increments instead of the feature are visible at all. Apply the fixes, rerun the gates, commit.
Consolidate the choices ledger, then close. When the last slice lands, the
choices.md you've been appending to per pass is build-order sediment: entries
banked early carry "provisional — revisit in slice N" verdicts that a later pass
silently resolved, entries a later pass reverted still sit there, and the same
choice may appear twice. The per-pass rule "banked is settled, never re-listed"
is what let that drift accumulate — so the final consolidation is its deliberate
exception. Before archiving, rewrite choices.md from scratch as the final
ledger: re-audit every banked choice against the final shipped code (not the
pass it landed in), collapse each provisional/needs-later entry to its actual
end state, drop anything a later pass superseded or reverted, and merge
duplicates. Keep it choices only — no gate results, e2e evidence, or
review-finding narration; those are reported elsewhere and are not decisions the
user now owns. Present it per audit-choices: grouped
by verdict, ranked least-confident-first, every entry ELI5 and standalone.
Then close the spec with close-spec.
A pass is done when code, spec handoff, verification evidence, review cleanup, and a focused commit all agree on the same current truth — then you start the next pass.
The spec is done — and only then is this skill done — when every slice and global TODO is closed, all gates are green, the whole-spec review has run and its fixes have landed, the handoff shows nothing left to pick up, and the spec has been archived with close-spec. Anything short of that is mid-implementation: keep going.
The final handback presents the choices ledger, not the diff, per audit-choices — a days-long unsupervised run earns its merge through this ledger; it is the user's review surface for everything decided without them. Hand over the consolidated ledger from step 11 (final state, verified against shipped code, choices only), never the raw per-pass append — a ledger still carrying "will be done in a later slice" verdicts tells the user you never went back to confirm it was.
Close the handback with the size of what you added — always last, after the ledger. The ledger says what was decided; this says what it cost. A short table over the whole run:
| added | deleted | net | |
|---|---|---|---|
| Production code (excl. comments) | |||
| Comments | |||
| Tests / harness | |||
| Specs & docs |
Then one paragraph naming the structural surfaces the run added — a new cron, table column, index, endpoint, config flag, dependency — because those are what the user now owns and maintains, and a line count alone hides them. Exclude formatter churn from files the run did not otherwise touch, and say so if you excluded any.
State the count plainly whatever it is. A large net addition for a small behavioural change is a finding to report, not a number to bury — and if you notice it here rather than at the pass that caused it, say which pass it was.
Slices are not the only unit of scope. A spec also records decisions — ledger rows, decision-table entries, invariants — and a decision can be agreed in planning but never turned into a slice, especially when it's orthogonal to the slices' theme. "All slices closed" then reads as done while that decision is silently unbuilt. Before declaring the spec done, reconcile every recorded decision against the shipped code, not just the slice list: each one is either implemented, or explicitly marked "no code needed." A decision with no owning slice is the classic silent miss — the close-spec audit is the backstop for it, not the first line of defense.
まだレビューはありません。使ってみた感想をお寄せください。
概要と使いどころ
Audit or rewrite AGENTS.md so it holds only lasting principles. Run it only when the user asks for it by name; never invoke it on your own.
日本語の概要は準備中です。原文の説明を表示しています。
Audit the choices an implementing agent made, not its diff — a pure decision audit that traces the session's history into a choices ledger, changes no code, and never blocks an unsupervised run. Working code still embeds architecture the user never chose; surface it because future work inherits it. Use when the user wants to review the decisions the AI made on their behalf, before merging or committing AI-implemented work, when integrating a delegated subagent's pass, or when a fix "works" but might be a point fix.
日本語の概要は準備中です。原文の説明を表示しています。
Audit and prioritize performance work through bounded-work and forward-progress checks. Use when asked to find performance issues, investigate freezes/500s/OOMs, review retries or polling for no-progress loops, rank a performance backlog, or implement the next simple performance fixes.
日本語の概要は準備中です。原文の説明を表示しています。
Audit whether tests earn their maintenance cost and which suite owns each contract. Use when pruning redundant tests, investigating implementation coupling or test-only production hooks, reviewing the value of proposed coverage, or auditing an entire subsystem. Use write-tests to implement the resulting test changes.
日本語の概要は準備中です。原文の説明を表示しています。
Run fast, progressive experiments to map tunable parameters and their effects. Use when optimizing code, prompts, configurations, or other artifacts against a goal or benchmark, or when a long research run needs a planned hypothesis queue, focused trials, and durable findings.
日本語の概要は準備中です。原文の説明を表示しています。
Use Claude Code as an independent `claude -p` subagent when the user explicitly asks for Claude, wants a second-agent opinion from Claude, or asks to delegate a well-scoped task to Claude. Supports selecting `--model` and thinking/effort level with defaults of `opus` and `high`.
日本語の概要は準備中です。原文の説明を表示しています。