Polish Code
Task Tracking
At the start of every invocation (including re-runs from Step 7), use TaskCreate to create a task for each step:
- Run
/stage skill
- Run
/run-checks skill
- Run
/review-code skill
- Run
/evaluate-findings skill
- Run
/apply-findings skill
- Run
/smoke-test skill
- Re-run
/polish-code skill if changed
Step 1: Run /stage Skill
Run the /stage skill.
Step 2: Run /run-checks Skill
Run the /run-checks skill.
Stage all changes made in this step before continuing.
Step 3: Run /review-code Skill
Run the /review-code skill on the staged changes. The diff command is git diff --cached.
Step 4: Run /evaluate-findings Skill
Run the /evaluate-findings skill on the results from Step 3.
Step 5: Run /apply-findings Skill
Record the staged tree with git write-tree before applying anything, and use TaskUpdate to add the tree ID it prints to this run's Step 7 task description.
Run the /apply-findings skill on the evaluated results.
When a defect this run fixes, in this step or in Step 6, is a further instance of a class of defect an earlier iteration already fixed, stop patching the individual instance and instead encode the root-cause invariant structurally — a shared guard or type, or a regression test that pins the class against the worked failures it must prevent. For a finding, make that the fix /apply-findings applies for it, so that skill's checks run on it. Recognize a class by its failure and what triggers it, so a further instance in another function or file counts as the class recurring. Count a defect that swings to the opposite failure after its fix, such as a check found too strict in one round and too lax in the next, as the same class recurring. In the same pass, audit the existing code against the newly encoded invariant and fix every instance it catches, including code written before it existed. When the recurring instance sits in code outside the changeset and so reaches the user as an escalated finding, make the invariant and that audit the remedy offered for it, with a fix to the individual instance as the narrower alternative. Treat recurrence on a new axis of the same invariant as a signal that the invariant is incomplete: widen it to cover the new axis rather than assuming the latest fix failed.
Record any fix whose remedy you deliberately narrowed as you make it, naming what the remedy covered and what it left, so Step 7 carries it forward without reconstructing the decision later.
When a fix ships with a regression test, confirm the test fails with the fix reverted, then restore the fix. When that test's outcome depends on the order of concurrent events, also run it many times over with the fix in place. Find what caused a failure on any of those runs, typically an ordering defect in the code or a missing synchronization point in the test, and fix it, then repeat both checks before this step closes.
Stage the fix immediately before mutating it (git add <file>), so git checkout -- <file> restores it exactly from the index. A file staged in an earlier step has an index copy older than the current edits, and restoring reverts them. Stage only the files about to be mutated: a broader restore point sweeps in working-tree changes the project may require stay uncommitted. When one also carries unrelated changes, write git diff <file> to a patch file, back the unrelated changes out of it (delete their + lines and turn their - lines into context lines), and stage it with git apply --cached --recount <patch>. Its index copy then lacks those changes, so copy that file aside before mutating it and restore it from the copy instead of the index. Each mutation edits the shared working tree in place, so hold anything that reads or builds that tree until the mutation is restored.
Before running the tests against a mutation, confirm it landed: git diff -- <file> shows the intended change against the index for each mutated file. An edit whose match pattern missed leaves the file untouched. Count a mutation as caught only when the test run reports a failing test, not when the mutation command or a step chained before the tests exited nonzero. When the mutation checks a particular test, count it only when that test is among those reported failing. When the suite crashes or hangs before that test reports, re-run the mutation against the narrowest test selection that includes that test: that run settles only whether that test catches the mutation, and a mutation it does not report on either stays uncounted.
When the fixed code combines several signals, also apply the plausible rewrites a maintainer might reach for — reordering the signals, substituting a fallback chain for a conjunction, dropping a term that looks redundant — and confirm each fails at least one test, then restore the fixed code. A rewrite that passes every test while changing behavior on some input means the tests pin the examples rather than the invariant; add the test that distinguishes it. When the fix guards against an unbounded loop or wait, bound the test itself so that reverting the fix fails rather than hangs: cap the iteration count for a loop; enforce a deadline for a wait. A deadline inside the test can depend on the test cooperating, so also run each such mutation under a timeout enforced from outside the test process.
When the fix changes when, whether, or how often a mechanism runs, mutate the changed line and run the whole affected suite, not only the test written for this fix: a fix can disarm tests that already existed, and those keep passing. Concentrate on the tests whose pass condition is an absence, and establish for each one that still passes whether it passes for the reason it did before. Those tests cannot distinguish a guard that rejected the work from a mechanism that never ran.
A test whose pass condition is an absence needs an assertion establishing the mechanism was reachable. Place that arming assertion before anything that can consume the state it reads. Placed after, the assertion holds whether or not the guard exists, and the test looks rigorous while pinning nothing.
A test that asserts only the direction of a numeric change passes on any movement of the right sign, including rounding noise or drift the fix did not cause. Assert the expected size of the change instead.
When the test still passes with the fix reverted, suspect the mutation before the test: confirm it reaches the branch under test and reproduces the original behavior rather than a third one. Name the branch under test and the original behavior it reproduces before running the mutation, then re-read the mutated lines. A mutation that lands a statement away from that branch also changes paths the fix never touched, and the resulting failure is indistinguishable from a caught mutation.
When the mutation is faithful and the test still passes, the test cannot observe the defect. Determine which of four shapes applies before reworking the test's setup:
- The assertion inspects output that is identical whether or not the defect is present. Assert on the mechanism itself rather than on the output it produces: teardown, cancellation, deduplication, and ran-only-once fixes leave no trace there.
- The inputs the test drives land the same way under the fixed value and the mutated one. When the fix seeds an initial value, drive an input on the far side of that seed; a test that only advances past it never observes the seed.
- An earlier guard against the same condition catches first, leaving the guard under test unreachable. Construct the ordering that reaches the later guard specifically. A defense-in-depth guard is reachable only through the window its predecessor does not cover, and an end-to-end exercise of the operation misses it systematically.
- The entry surface the test drives is itself closed for the window the guard covers, so no input the test can send reaches the guard. Where the earlier shape puts a predecessor in the code path, this one puts the block ahead of it: when the fix guards a repeat invocation, a busy or pending state on that affordance refuses the repeat before the guard sees it, and the test passes for the wrong reason. Drive an entry point that carries no such state instead.
A test that genuinely cannot be made to fail does not pin the behavior; say so rather than counting it as coverage.
After every mutation in this step, re-run whatever that mutation was checked against and confirm it passes again and the test command itself exits 0, not a filter piped after it, before reporting the result. A clean git status looks identical whether the fix was restored or deleted. A run that a mutation made fail can leave behind what a passing run cleans up, such as temporary files and spawned processes. Before that re-run, clear what the failed run left, identified by what the run itself started or named. Leave anything whose origin that does not establish, and name it when reporting the result.
A project command run to verify a fix writes to the shared tree the same way a mutation does. Establish whether it writes tracked files before running it, and read git status --short afterward: revert what it wrote, so files it regenerated are not swept into the changeset by the staging below.
Stage all changes made in this step before continuing.
Step 6: Run /smoke-test Skill
Capture git status --short, git diff --cached | git hash-object --stdin, git diff | git hash-object --stdin, and git symbolic-ref --short -q HEAD before spawning.
When this run was supplied a smoke baseline, compare it against the first three outputs and git rev-parse HEAD. When all four match, report the result recorded with that baseline as carried forward and close this step, running neither the /smoke-test skill nor a test run. On any difference, or with no baseline supplied, continue.
Run the /smoke-test skill to produce the smoke test plan.
Delegate test execution to an Agent tool call (model: "opus", no name). Wait for it to report before continuing; do not relaunch it if it has not yet reported. Pass the plan and the diff command (git diff --cached) to the subagent, and instruct it to invoke /test-run-rules via the Skill tool before executing the plan. State in its prompt that the writes the plan's Setup contract authorizes are already approved, and that any write outside that enumeration leaves its scenario blocked.
Verify the tree: re-run all four commands when the subagent returns, including when it terminates early or reports incomplete results. Delete what the subagent created, revert what it modified or staged, and return HEAD to the captured branch, leaving everything the pre-spawn capture already showed untouched.
If any test fails, or the subagent reports a defect in the staged changes outside the plan's scenarios and the code confirms it, fix the issues and stage the fixes. When every planned test passed, this step made no fix, and the verification above found nothing to delete or revert, use TaskUpdate to add the smoke baseline to this run's Step 7 task description: the first three captured outputs, git rev-parse HEAD, and the result the run reported.
Step 7: Re-run /polish-code Skill if Changed
Check whether any file was edited during Steps 5-6: git diff --name-only <tree> --cached lists them, where <tree> is the tree ID this run's Step 5 recorded in this step's task description. Any edit counts.
The iteration number below refers to the /polish-code run currently executing Step 7. It is not the iteration number of a prospective re-run. Iteration 1 is the initial run; iteration 2 is the first auto-re-run; iteration 3 is the second auto-re-run; iteration 4 and beyond exist only when the user opts in at the hard-cap ask. Iterations 1 and 2 always follow the classification gate (they never trigger the hard cap at their own Step 7, even when the auto-re-run they spawn would be iteration 3). The hard cap fires at the end of iteration 3 and every iteration thereafter.
Iterations 1 and 2, if changes were made, classify what Steps 5-6 edited:
- Structural edits (fixed bugs, new or removed functions, changed function signatures, moved code between files, changed control flow, added or removed dependencies, corrected a stale or wrong comment that was itself a documentation bug) — run
/polish-code again using the Skill tool. Scope the diff command to what changed since this iteration's review: use git diff <tree> --cached as the diff command for /review-code. Smoke test scope remains unchanged (full feature scope, not narrowed to that diff). If the round contains both structural and in-place edits, treat it as structural and re-run automatically.
- In-place edits only (renamed local variables without changing behavior, reformatted, adjusted whitespace, edited neutral comments) — output a summary of what changed, then use
AskUserQuestion to ask whether to run one more round or stop here. Do not silently continue or silently stop.
Iterations 1 and 2, if changes were made but you believe re-running is unnecessary, use AskUserQuestion to ask for skip permission. Do not skip silently.
Judge whether another round is worthwhile by the trend across iterations: when rounds have stopped surfacing defects (wrong behavior, security exposures, broken contracts) and keep surfacing improvements of kinds earlier rounds already applied, the loop has converged even though the edits were structural — recommend stopping. A round that surfaces no defects is the termination signal; never add a confirmation round, an extra reviewer, or review steps beyond this skill's own.
Treat reversal as the stronger signal: when a round's accepted findings undo an earlier round's accepted findings on the same lines, the reviewers are trading equally defensible positions rather than converging on one answer. Recommend stopping there even though such edits classify as structural. Keep the current round's version of the reversed code, which is as likely to be the better answer as the one it replaced.
Weigh where a round's defects live alongside whether it found any. When the change the run set out to make has been stable for two or more rounds and every new defect sits in verification scaffolding an earlier round added — a probe, a gate, a check, or the documentation describing them — the loop is generating its own work rather than converging on the change. Recommend stopping there even though the round surfaced defects.
Wherever this step asks or recommends stopping rather than re-running automatically, weigh the files Steps 5-6 changed that this iteration's review did not read: every iteration reviews the diff as it stood before its own fixes, so those files have had no review pass. Let them argue for one more round whose /review-code diff command is git diff <tree> --cached. The reversal and defect-location signals outrank that, so a round tripping either one stops even with files left unread. Files a later round's review reads stop counting toward this.
Every stop this step reaches with changes pending names those unread files, by path, or by cluster with a count when there are many. State alongside them that stopping leaves those files covered by the Step 6 smoke run and by Step 5's per-fix verification, and not by /review-code or by the Step 2 check gate, which ran before they existed. Name separately any file a Step 6 fix edited: the smoke run and Step 5's verification preceded that fix as well.
Iteration 3 or later, if Steps 5-6 of this run made changes, the hard cap is reached. This replaces the classification gate above for iteration 3 and every iteration after it. Output a summary of what is still changing, whether it is structural or in-place, and where this round's defects lived (in the product, or in the build, CI, and gate scaffolding around it). Then use AskUserQuestion to offer three options: run another iteration, whose review scope includes the unread files named above; stop iterating, accepting the current state so the workflow moves on; or escalate to /consult-oracle for a different perspective on the remaining issues.
Before acting on any gate above, check for a context signal. Call the mcp__context-level__read tool: a reading of 25% or less remaining is one. A request from the user to compact, arriving since the session last compacted, is the other. When either is present, turn whatever this step reached, including an automatic re-run and a run that edited nothing, into one AskUserQuestion call. Its first question asks whether to compact before what comes next: "Handoff, then compact", marked recommended, or keep going in this session, which acts on the gate's answer here. Leave it out when the user asked to compact. Its second question offers the choices of the gate this step reached, with the recommendation that gate would make; in place of an automatic re-run, it offers run another iteration, marked recommended, or stop iterating. Leave it out when no file was edited in Steps 5-6. Make no AskUserQuestion call when both are left out. When no file was edited and TaskList shows no pending task that no /polish-code iteration created, nothing remains to resume: ignore the signal and let the loop end.
On "Handoff, then compact", or when the user asked to compact, check first whether the gate's answer is stop iterating while TaskList shows no pending task that no /polish-code iteration created. Then nothing remains to resume: tell the user so and close this step without a handoff. Otherwise run the /create-handoff skill, or, when this session already wrote a handoff, edit that file. Point its next step at what the gate's answer leads to, and once the handoff is written, update this step's task as that branch says:
- Run another iteration — the re-run. The handoff carries the number of the iteration about to run, the already-adjudicated list this step supplies to it, the smoke baseline when Step 6 recorded one, and its
/review-code diff command. Use TaskUpdate to set this step's task description to the re-run: run /polish-code as that iteration with that diff command, reading the already-adjudicated list and the smoke baseline from the handoff at its path. Leave the task in progress.
- Escalate to
/consult-oracle — the escalation. The handoff carries the summary above and the same state as on another iteration. Use TaskUpdate to set this step's task description to that escalation, reading that state from the handoff at its path. Leave the task in progress.
- Stop iterating, or no file was edited — the first pending task that no
/polish-code iteration created. The handoff carries the unread files this stop leaves, and the message ending the turn names them as every stop does. Mark this step's task completed, along with every Step 7 task earlier iterations left in progress.
Then end the turn in place of the TaskList call that closes this step, telling the user to run /compact and then reply "continue".
The re-invocation is a full, fresh run of this skill. Every step (1-7) executes with its own task tracking and skill invocations, apart from a smoke result Step 6 carries forward. The narrowed diff command only affects what /review-code reads. It does not affect which steps run or whether skills are invoked. Whichever gate above sends the run into another iteration, supply that iteration with every Skip and Escalate verdict recorded so far, and every Apply whose remedy Step 5 recorded as narrowed, across this run and earlier iterations, as the already-adjudicated list for /review-code, one line each: the finding, its verdict, and the recorded reason. For an Escalate the user resolved, the reason attributes to the user only what they decided; the details you picked while implementing it stay open to review. A narrowed Apply carries what the remedy covered and what it left, so the untouched remainder reads as settled rather than as an unaddressed gap. Fresh task tracking leaves that list intact. A finding that re-proposes a remedy an earlier round narrowed stays in scope regardless of the list: the remainder having since caused a defect is evidence the earlier reason did not account for, and it is the signal Step 5's recurring-class rule depends on. Supply the next iteration as well with the smoke baseline this run's Step 6 recorded, when it recorded one.
Then use the TaskList tool and proceed to any remaining task.
Rules
- Every step must run in every iteration.
/review-code covers correctness, security, consistency, API usage, coverage, and simplicity across parallel internal reviewers plus peer review. /evaluate-findings is a judgment gate that must run before /apply-findings.
- Each step must invoke its designated skill via the Skill tool, not be replaced by inline reasoning or agent calls.
- Re-invocations from Step 7 are full runs with fresh task tracking and complete skill invocations.
- Step 6 carrying a smoke result forward on a matching baseline is the one exception to the rules above.