subagent-driven-development
Structured-artifact format boundary
New full-lane workflows default to HTML through the installed engine. Existing
HTML pairs use the same path. Load
{{PLUGIN_ROOT}}/includes/artifacts.md before HTML artifact I/O;
its typed identity, binding, receipt, contextual-read, and scope-state rules
replace rendered-document parsing and direct status-cell edits. Existing
authoritative Markdown pairs stay on their exact legacy path without conversion.
A missing, invalid, or unqualified HTML capability must fail closed rather than
fall back to new Markdown authority. HTML --farm dispatch remains disabled.
Keep every other workflow gate, including human checkpoints, unchanged.
One task, one fresh subagent, two reviews, proof on a fresh run. Routed to by /sprint (full plan,
autonomous) and by executing-plans (one batch at a time, with human checkpoints between batches).
The loop processes tasks in dependency order and never trusts a self-report.
Pre-flight
Read these, or STOP and surface the gap — never guess scope, command, or obligation:
{{PROJECT_DIR}}/.codearbiter/CONTEXT.md — the stage: frontmatter (the maturity value) and project context.
- The discovered authoritative spec/plan pair. Existing Markdown pairs retain
their current ledger. For HTML, load
{{PLUGIN_ROOT}}/includes/artifacts.md,
retain the exact artifact identities and engine-written receipts, and read
the approved spec binding, obligations, paths, verification definitions,
dependencies, and state only through identity, index, outline,
symbol-scoped read, and eligible. The engine eligible response is authoritative;
the workflow must not infer authority from rendered text, checkboxes, status
words, digests, or caller-supplied labels, and must not read or write a shadow
Markdown ledger.
{{PROJECT_DIR}}/.codearbiter/tech-stack.md — build, test, and verification invocations; file layout; the scope-to-author mapping.
{{PROJECT_DIR}}/.codearbiter/security-controls.md — only when a task touches a security boundary (auth, crypto, secrets, a trust boundary).
Optional scope parameter: when invoked by executing-plans, a list of task IDs is passed. The
loop processes only those tasks (in their internal dependency order). When scope is absent (the
/sprint path), the loop processes the full plan from first unblocked task to last.
Phase 1 — Task selection · gate: BLOCK
Pull the next unblocked task from the plan in dependency order. When a scope was passed, restrict
selection to tasks in that list. For Markdown, select only a task whose status is PENDING and whose
dependencies are ACCEPTED; BLOCKED tasks are not selectable until the owning plan workflow
records a supported transition. For HTML, accept only tasks returned by eligible for the exact
selected plan identity. A task is one verifiable unit of work with a path set, a spec obligation,
and a verification command.
For HTML, select only a task returned as eligible for the exact selected plan
identity. Follow contextual read pages to completion, retain their context
ticket, and use task-start before author dispatch. Do not translate a displayed
status or a copied task label into a transition.
After process recreation, rediscover the same authoritative pair and query
identity and eligible again. Reconcile an interrupted IN_PROGRESS task only
through an authority-backed task-reconcile, then redispatch with task-start
and a fresh complete context ticket. A task in REVIEW is not accepted on
resume; obtain fresh evidence when required and finish the current combined
scope review before accept-scope.
- For Markdown, confirm every dependency task is
ACCEPTED before selecting.
- For HTML, use the engine's eligibility and current evidence; do not replace its same-checkpoint
REVIEW rule
with the legacy ACCEPTED sentence or infer eligibility from a displayed status.
- Confirm no unresolved
[CONFIRM-NN] blocks the task. One that does halts the loop — see Hard rules.
Gate: exactly one task selected, dependency-clean, with its spec obligation and verification command in hand.
Phase 2 — Implementation dispatch · gate: BLOCK
Farm path (when <slug>.plan.json exists alongside the .md plan): skip the subagent dispatch
loop below and follow {{PLUGIN_ROOT}}/skills/subagent-driven-development/references/farm-dispatch.md.
The farm path replaces only the authoring step for the plan's tasks (cheap Zen workers under hard
gates instead of premium subagents); it does not replace review — every task the farm reports green
is still routed through Phases 3–5 before acceptance. The cost arbitrage is in who writes the code,
never in whether it is reviewed. In brief: select a model (canary-probe with a cache→websearch
fallback ladder), {{IF:pi}}invoke the trusted codearbiter_farm_preview tool with the project-relative
plan path{{ELSE}}dispatch {{PLUGIN_ROOT}}/tools/farm.js{{END}}, honor a circuit-breaker abort as a hard-gate STOP, then for
each result either accept-after-Phases-3–5 (green) or re-dispatch via premium Phase 2 (escalate).
Results stream to .farm/farm-results.jsonl and are consumed in completion order — Phase 3 + Phase 5
per green task as it lands, Phase 4 still the once-per-scope barrier (reconcile against farm-report.json
on abort). The reference has the full step-by-step.
Normal path (no plan.json): dispatch ONE fresh subagent for the selected task — backend-author, frontend-author, or
infra-author by the scope mapping in tech-stack.md
({{PLUGIN_ROOT}}/agents/<name>.md). A fresh context per task is the whole point: no carried-over
assumptions, no accumulated drift.
The subagent works test-first by routing through the tdd skill ({{PLUGIN_ROOT}}/skills/tdd/SKILL.md) — no implementation code before
tdd Phase 1. Brief it with the task's path set, its spec obligation, and its verification command.
Nothing else from prior tasks leaks in.
Before dispatch, obtain current bounded inputs for the selected task paths:
existing code-map text if available, assessed provenance, scoped command
records and collector observations, applicable constraint references, and
host-effective native instructions in precedence order. Call the private
_artifactpromptlib.compose_feature_actor_input with the chosen author name,
brief, exact approved spec and plan IDs, task ID, task paths, and those inputs.
Send its returned input as the fresh child's actual input. The returned
packet never grants task start, source-snapshot admission, or permission to
mutate or execute a discovered command. Resolve an applicable critical
unresolved constraint before that material action. An unavailable optional
map permits bounded read-only orientation, not an invented complete map.
For each author or reviewer dispatch, call _contextselectlib.prepare_actor_delivery
on that actor's freshly selected packet with the actual actor/task, absolute
worktree, current source identity, and host epoch when observable. Send its
bounded delivery_text to that child on deliver and retain the counted
attempt. Treat observed only as caller-reported load/input evidence; do not
use a shared session or persona/read marker as a receipt. On blocked, stop
the material action and report the delivery gap. Re-select after resume,
compaction, scope expansion, or worktree switch; missing epochs get at most
two counted deliveries for the same binding.
Gate: all six tdd phases must be green before acceptance. Never redispatch merely to evade a failing gate.
Under an approved /sprint, an ordinary implementation/test failure routes to Recovery within the approved sprint in {{PLUGIN_ROOT}}/SPRINT.md; rerun the original gate after an evidence-led correction.
An attended invocation returns the blocked prerequisite to its caller. A real authority or security block
halts the affected work and is surfaced under the hard rules, in either mode.
Phase 3 — Spec-compliance review · gate: BLOCK
Did the change satisfy the task's obligation? Measure the result against the spec line the task
traces to — not against whether tests merely pass.
For an independent spec reviewer, freshly compose the same task-scoped input
with actor spec-reviewer and the exact artifact/task IDs, then send the
returned input to that reviewer before its substantive verdict. Do not
substitute the author's packet or the coordinator's summary for reviewer
delivery; the reviewer independently checks source identity and constraints.
- Every acceptance claim in the task's obligation is met by the change.
- Scope is clean: nothing implemented beyond the task; nothing required by it omitted.
- Out-of-scope work the subagent noticed is recorded with an inline
[NEEDS-TRIAGE] marker — never
acted on inside this task.
For HTML, capture the actual spec-review event against the exact task and
artifact identities. After the governed verification runner has produced its
separate receipt, use task-review with those two engine-written receipts. A
digest or the reviewer/subagent's label is not evidence by itself.
Gate: the obligation is fully satisfied and scope is clean. A shortfall returns the task to Phase 2
with a corrective brief.
Phase 4 — Quality review (once per scope) · gate: BLOCK
Runs ONCE per scope — after every task in the current scope has cleared Phase 3 and Phase 5 — over
the combined diff of the scope, not per 2–5-minute task. Per-task review at that granularity
costs more context than the work and catches nothing the batch diff doesn't; the batch boundary is
where review pays. (A scope of one task reviews that task's diff — same rule, degenerate case.)
Dispatch the reviewers applicable to what the combined diff touches, then finding-triage
({{PLUGIN_ROOT}}/agents/finding-triage.md) to classify every finding by severity. Select
reviewers by the diff, not blanket — dispatching an irrelevant reviewer wastes a context:
For each dispatched reviewer, compose and send a fresh actor input with the
union of paths in the combined scope, the current bounded context inputs,
exact approved spec/plan IDs, and the scope task binding. Preserve independent
review judgment and all existing review gates. A packet read only by the
coordinator is not reviewer delivery; unresolved critical constraints block a
substantive verdict while bounded source inspection remains possible.
security-reviewer ({{PLUGIN_ROOT}}/agents/security-reviewer.md) — any security-relevant path (authn/authz, deploy, CI, trust boundary).
auth-crypto-reviewer ({{PLUGIN_ROOT}}/agents/auth-crypto-reviewer.md) — auth, crypto, key, or secret changes.
dependency-reviewer ({{PLUGIN_ROOT}}/agents/dependency-reviewer.md) — package.json / lockfile / base-image changes.
migration-reviewer ({{PLUGIN_ROOT}}/agents/migration-reviewer.md) — DB migration add/modify.
(Do NOT dispatch grader or scout — they are INTERNAL to decision-variance and must never be
dispatched here.) If the change touches none of the above domains, the quality bar is tdd's own gates
plus coverage-auditor (already run in tdd Phase 4) — record that and proceed.
- A security CRITICAL finding halts the loop — see Hard rules.
- A HIGH finding returns the offending task(s) — attributed by file — to Phase 2; the scope's
quality review re-runs over the corrected combined diff.
- MEDIUM and LOW findings are recorded; the active caller owns their disposition. Under an approved
/sprint, use its existing delegated decision rules rather than adding a per-finding user checkpoint.
In attended execution the user retains that decision; mandatory project policy still applies.
Gate: no CRITICAL, no HIGH across the scope's combined diff. Nothing in the scope is ACCEPTED
until this passes.
Phase 5 — Verification · gate: BLOCK
Verification-before-completion: apply the shared fresh-run discipline in
{{PLUGIN_ROOT}}/includes/fresh-verification.md, with the task's verification command from the
plan as the target. Run it yourself in a clean invocation; do not accept a logged result from Phase 2.
A non-zero exit, or output that does not demonstrate the obligation, returns the task to Phase 2.
Gate: the verification command exits clean and its output demonstrates the obligation. Only then.
Phase 6 — Accept and advance · gate: BLOCK
Mark a task accepted only when its spec-compliance review and fresh verification
passed AND the scope's Phase 4 quality review passed. For Markdown, record the
existing status-cell transition in the plan. For HTML, capture the actual
combined quality-review event and call accept-scope once for the whole scope;
the resulting engine state and acceptance receipts are the resume ledger. Never
edit an HTML status cell or create a Markdown counterpart. In either format, an
acceptance that lives only in conversation context is lost to interruption.
- Tasks remain in the current scope → return to Phase 1.
- Scoped invocation (
scope was passed by executing-plans): all scoped tasks ACCEPTED → signal
batch complete and return to executing-plans. Do NOT hand to commit-gate; the caller owns that decision.
- Full-plan invocation (no
scope, i.e. /sprint): plan complete → hand the branch to
commit-gate, then to the caller's finishing step. The loop does not commit on its own authority.
When either path claims the entire HTML plan is currently accepted, invoke the
installed bridge's _preflight_current_acceptance with the exact
_resolve_workflow_pair result, repository-bound client, and engine-returned
spec/plan IDs. Do not pass or derive authority from a caller label, digest,
status, or receipt. A refusal means the plan is not complete; return to the
appropriate execution/reconciliation phase instead of handing off.
Gate: every task in the current scope ACCEPTED, the suite green, ready for the caller's next step.
Hard rules
- MUST dispatch a fresh subagent per task — never reuse one context across tasks.
- MUST NOT write implementation code before the task's
tdd Phase 1 completes.
- MUST accept a task only when both reviews pass AND verification passes on a fresh run.
- MUST NOT accept a task on a subagent's self-report — run the verification command and read its exit code.
- MUST halt and surface a real authority or security block, including security CRITICAL findings and
unresolved
[CONFIRM-NN] decisions. Ordinary TDD/verification failures and finish-time stale proof
under an approved sprint take Recovery within the approved sprint in {{PLUGIN_ROOT}}/SPRINT.md.
A failed gate never auto-passes, and recovery never grants missing commit authority.
- MUST NOT commit — hand the accepted branch to
commit-gate.
- MUST mark out-of-scope findings with an inline
[NEEDS-TRIAGE] marker and never act on them inside the task.
- MUST NOT invoke
farm.js before writing meta.model into plan.json (or setting FARM_MODEL) — the dispatcher fails loudly otherwise.
- MUST NOT skip model selection (Step 1) when
FARM_MODEL is not set — exhaust the canary→cache→websearch fallback ladder before BLOCKing; never blind-invoke with an unknown model id.
- MUST route every green farm task through Phases 3–5 before acceptance — the farm replaces authoring, never review. A cheap model gets the same scrutiny as a premium subagent, not less.
- MUST treat a single task's drift/gaming/tampered-test escalation as model incapacity (re-dispatch via premium Phase 2); raise
[CONFIRM-NN] and HALT only when multiple tasks drift onto the same out-of-scope path (a real spec gap).
- MUST treat a
farm.js circuit-breaker abort (aborted: true) as a hard-gate STOP — surface to the user; do not silently re-dispatch the whole slice to premium.
- MUST NOT dispatch
grader or scout in Phase 4 — they are INTERNAL to decision-variance.