Use when reviewing UI for accessibility — WCAG 2.2 AA, keyboard nav, focus, ARIA, contrast, screen-reader semantics — even on 'is this a11y-OK?' or 'mach das barrierefrei'.
日本語の概要は準備中です。原文の説明を表示しています。
Use when orchestrating implementer/judge subagents — form gate + nine modes (do-and-judge ±two-stage, steps/parallel/worktrees, competitively, debate, live-app-judge, adversarial-council).
インストールする前に、エージェントに与えられる指示の中身を確認できます。
Do NOT use when:
token-optimizer index branchLand a verified change (or set of changes) by combining implementer
and judge subagents in a mode chosen deliberately, with model pairing
read from .agent-settings.yml — never silently improvised.
NO JUDGE ON THE SAME MODEL AS THE IMPLEMENTER ON THE SAME CONTEXT.
Same model + same context = same blind spots. The whole point of a
judge is a fresh pair of eyes. If .agent-settings.yml resolves to
identical implementer and judge models, surface the mismatch before
running — do not silently continue.
Within the Reasoning Discipline Protocol, dispatch independent subtasks to
parallel subagents by default and keep working while they run (async), rather
than blocking on each return — intervene only if one goes off track. Engage per
rdp-gate.
"By default" is governed by the delegation-policy
rule — the single source of the auto-trigger. It gates on the activation context
(auto-orchestration-activation):
dispatch only when emergency.orchestration_halt is not set, the host
manifest reports subagent_spawn: true, and the task is classified delegable.
A matched signal → dispatch, surfacing mode + per-subtask tiers in one line;
an ambiguous verdict → ask, always; any gate failing → in-session no-op.
There is no more per-layer on/off setting (always-on orchestration). Never
lifts a safety floor.
Every dispatched worker prompt obeys five rules that prevent the two classic
handoff failures (lossy re-summarization dropping the user's requirements;
over-scripted prompts that break on first contingency): (a) user
constraints verbatim, (b) describe the goal — don't script the approach,
(c) translate environment paths into the worker's sandbox, (d) pre-declare
check-in conditions, (e) attach relevant knowledge read-only (auto-surface,
never auto-write — ADR-098 floor). The five rules verbatim:
subagent-spawn-contract § Worker-prompt rules.
When to delegate at all is delegation-policy;
the spawn boundary is the subagent-spawn-contract.
Ordered / fan-out hand-offs embed each step's return verbatim in the next prompt and state what to do with it (never "continue from before" — the lossy re-summarization failure). Two worked shapes: subagent-spawn-contract § Hand-off worked examples.
With auto-dispatch on by default (ADR-117), mode selection happens without a human in the loop — so the FORM is decided by a static table first, and only then is the specific mode picked inside that form. Static table only: no learned routing, no self-modifying selector (rejected, stays rejected).
| Task shape (structural signal) | Form | Modes in the form |
|---|---|---|
| ≥ 2 independent, verifiable slices | parallel | do-in-parallel, do-competitively |
| Multi-step cross-wing chain needing filesystem isolation | worktrees | do-in-worktrees |
| Ordered steps with declared dependencies | steps | do-in-steps |
| Single change with non-trivial risk / contested spec / decision | judge | do-and-judge, do-and-judge-two-stage, judge-with-debate, do-with-live-app-judge |
| High-risk change needing defect-FINDING coverage (opt-in, advisory) | verify-council | adversarial-verification-council (default-off; subagents.adversarial_council) |
| Single slice below the delegability floor, unstructured, or frontier-priced | none | no dispatch — run in-session |
Rules:
auto_dispatch.ts::classifyTask (slice count, dependency declarations,
size floor) — it never re-interprets the task text on vibes.none (in-session), never a speculative spawn — the
delegation-policy default.do-in-parallel-shaped fan-out from one parent has this shape by
construction (a known, unfixable upstream cost — see
prompts/README.md § Prompt-cache discipline). NEVER
encode "always fork" — that was proposed and cut (council 2026-07-30):
the cache-sharing benefit cannot be predicted before the fork happens.dispatch_mode field, mode id
or none) so the gate's value is measurable inside the ADR-117
prove-or-drop window.Incident-style severity tiers (Critical / High / Medium / Low) refine
composition and activation within the form the static gate already
picked — severity never overrides the form gate, the Iron Law, or any
safety floor. Guidance, not a new object class (persona-catalog
disposition). Severity→composition table + escalation rule →
subagent-modes-detail § Severity-conditioned team composition.
Each mode has a decision row: when to use, when not, and the expected
model pairing. Defaults come from
subagent-configuration.
Descriptive lookup material (per-mode topology table, anti-drift default,
glossary) lives in
subagent-topologies —
pull it for capacity planning; it is metadata, not runtime-enforced.
Per-mode decision rows (when to use / when not / model pairing) and the
mode-2 stage-routing contract live in
subagent-modes-detail § Modes 1–6 —
pull them at dispatch time.
Implementer produces a diff; judge reviews; loop applies, revises, or hands off. Hard ceiling: two revision cycles, then stop and hand back to the user.
Implementer produces a diff; two judges run sequentially — spec
compliance first, code quality second. Stage-one BLOCKED shortcuts the
loop (no point quality-reviewing a diff that misses the spec). Stage
routing + why-two-stages rationale → modes-detail § Mode 2.
Plan is split into N steps; judge runs between steps. A step that fails judgment is revised before the next step starts. Used for multi-file changes where a mid-plan mistake would cascade.
Independent slices run concurrently. No judge per slice — judge runs
once on the aggregated result. Parallelism capped by
subagents.max_parallel in .agent-settings.yml.
Multiple implementers produce candidate diffs for the same slice. Judge picks the winner and rejects the losers. Expensive — use only when the solution space is genuinely broad.
Two judges each produce a verdict; a meta-judge reconciles disagreements. Used for high-stakes changes (security, data migration, public API) where a single judge is too easy to fool.
Mode 6 = go/no-go (strict-er verdict wins); for defect-FINDING coverage (the union of what diverse models catch) use Mode 9.
Cross-wing or cross-skill chain executed across isolated git worktrees —
each handoff runs in its own worktree so one step's workspace state never
leaks into the next. Use for a multi-step cross-wing chain (≥2 senior
skills, each ≥30 min); not for fast iteration under 30 min (overhead
dominates). Full handoff shape, example chain, competitive per-candidate
isolation, and the no-auto-merge Hard Floor →
subagent-modes-detail § Mode 7.
Implementer ships the change AND starts the dev server; the judge drives
the RUNNING app (Playwright / browser) against a written rubric, never
reading the diff. Use for UI-heavy change where "looks right in the diff"
≠ "works in the app"; not for backend/logic (a diff judge is cheaper).
Experimental until verdict_changed_outcome telemetry proves it. Rubric,
adoption gate, and the async-verifier future candidate →
subagent-modes-detail § Mode 8.
A panel of N (default 2) distinct-model skeptics red-teams a real,
already-verified change through the judge-* lenses — each prompted to break
it. Returns reconcile deterministically
(_lib/adversarial_reconcile.ts)
into one findings-by-severity envelope with provenance + cross-model confidence
(schemas/adversarial-findings.json). Unlike
Mode 6 it emits a findings-union, not a go/no-go verdict. Advisory only —
never auto-gates (Hard Floor). Default-off (subagents.adversarial_council);
opt-in high-risk changes only; registered claim + high-risk tier need
cross-vendor skeptics. Invariants, skeptic prompt, reconciliation, prove-or-drop
gate → subagent-modes-detail § Mode 9
Disposition (2026-07-28, honest null). The adversarial-council finding- coverage benchmark resolved as a published null (see
docs/benchmark.md) — this mode is NOT sold as a defect-detection capability. It stays default-off bound to that null; scheduled for removal at the next major unless external evidence (a consumer-filed case where the panel surfaced a real defect the single verifier missed) appears first. Its remaining honest value is perspective diversity + decision documentation, nothing more. (ADR-122).
Every implementer or judge return must conform to
schemas/subagent-status.json. Exactly
four statuses — DONE · DONE_WITH_CONCERNS · NEEDS_CONTEXT ·
BLOCKED — no free-form alternatives; orchestrators route on status
mechanically. Meaning/required-keys table, why-fixed rationale, and the
NEEDS_CONTEXT-vs-BLOCKED distinction →
subagent-modes-detail § Status taxonomy.
The envelope is the ONLY return channel (token-economy-dispatch
Phase 6). A worker writes its full output to disk (runtime artifact dir,
gitignored) and returns the bounded envelope — summary +
artifact_paths + verdict, size caps validator-enforced
(_lib/subagent_response.ts: summary ≤ 2,000 chars, whole envelope
≤ 12,000). The orchestrator reads FROM the artifact paths on demand —
never instructs a worker to paste its full result into the return, and
never ingests a transcript-shaped return wholesale: dispatching N workers
grows the orchestrator context by N envelopes, not N transcripts.
Each mode's literal dispatch template lives under
prompts/{mode}.md. The orchestrator loads the
matching prompt at dispatch time and substitutes {{placeholders}}.
Edits to a prompt do not bloat this skill against the 400-line sunset
trigger. Eight prompt files cover modes 1–7 and 9 (the standalone judge
reuses prompts/do-and-judge.md; mode 8's
live-app rubric lives in
subagent-modes-detail § Mode 8);
each prompt cites all four taxonomy statuses — see
prompts/README.md.
Before picking a mode, check:
Do not pick a mode until these four questions have concrete answers.
Read .agent-settings.yml:
subagents.implementer_model → empty = session modelsubagents.judge_model → empty = one tier above implementersubagents.max_parallel → integer, default 3If resolution produces an unknown alias or implementer == judge in the same context, stop and report. Do not improvise.
Run the form gate first, then match task shape to one of the nine modes. When two modes could fit,
prefer the cheaper one (do-and-judge < do-and-judge-two-stage <
do-in-steps < do-in-parallel < do-competitively <
judge-with-debate < do-in-worktrees).
Mode 6 (do-in-worktrees) is instruction-only (ADR-229). There is no
setting to resolve — worktrees.mode was deleted:
| Situation | Mode 6 |
|---|---|
| The user asked for a worktree chain in the chat ("do this in a worktree", "use mode 6") | Eligible. Proceed via using-git-worktrees; the request is the permission, so there is no further ask. |
| Anything else | Not eligible, however well the chain shape fits. Fall back to mode 3 (do-in-steps) — the same step-by-step chain, in place on the current branch. Do not offer mode 6 as an option. |
Because mode 6 is unreachable unprompted, it is effectively off the cheapness ladder above unless the user named it; read that ordering as covering modes 1–8 in the default case.
Use the matching dispatch prompt and orchestrate inline via this skill. Describe each dispatch step explicitly in chat so the user can follow it.
do-and-judge → prompts/do-and-judge.mddo-and-judge-two-stage → prompts/do-and-judge-two-stage.mddo-in-steps → prompts/do-in-steps.mddo-in-parallel → prompts/do-in-parallel.mddo-competitively → prompts/do-competitively.mdjudge-with-debate → prompts/judge-with-debate.mddo-in-worktrees → prompts/do-in-worktrees.mdadversarial-verification-council → prompts/adversarial-verification-council.mdprompts/do-and-judge.mdFollow the output format below. Never merge a diff without reporting the judge verdict.
After every auto-dispatched run, write one audit-log-v1 line
(input_kind: "orchestration") via the orchestration_record recorder —
never hand-author the JSON. Line shape, field semantics, recorder
invocation, and token_delta sourcing priority →
orchestration-telemetry § Emit procedure.
Skip emit when emergency.orchestration_halt is set or spawn_count == 0
(in-session run).
do-and-judge; then hand back to the user.do-in-parallel on overlapping slices — race conditions,
conflicting diffs. Verify independence before splitting.do-competitively — N implementers + 1 judge =
N+1 subagent calls for one slice. Confirm budget before dispatch.do-and-judge without user
consentdo-in-parallel on slices that touch shared files| Task | Skill / command |
|---|---|
| Configuration reference | subagent-configuration |
| Do-and-judge loop | Inline — see prompts/do-and-judge.md |
| Stepwise plan with judge gates | Inline — see prompts/do-in-steps.md |
| Standalone judge on an existing diff | Inline — see judge prompt in prompts/do-and-judge.md |
| External / networked second opinion | ai-council |
| Cross-model review WITH repo access | /team (collaborative; subagents are in-session same-weights) |
| Verifying completeness | verify-completion-evidence |
| What a subagent owns vs never owns | subagent-boundary |
まだレビューはありません。使ってみた感想をお寄せください。
概要と使いどころ
Use when reviewing UI for accessibility — WCAG 2.2 AA, keyboard nav, focus, ARIA, contrast, screen-reader semantics — even on 'is this a11y-OK?' or 'mach das barrierefrei'.
日本語の概要は準備中です。原文の説明を表示しています。
Use when defining or auditing the activation event — aha-moment selection, retention correlation, falsifiable definition. Triggers on 'what is our aha moment', 'redefine activation'.
日本語の概要は準備中です。原文の説明を表示しています。
Use when capturing an architectural decision — file naming, next ADR number, Status / Context / Decision / Consequences, index regen; fires even without saying 'ADR'.
日本語の概要は準備中です。原文の説明を表示しています。
Adversarial critique — devil's advocate, stress-test, honest teardown ('poke holes', 'be brutal', 'was hältst du davon'); explicit request only. Routine code or design review → code-review.
日本語の概要は準備中です。原文の説明を表示しています。
Use when reading, creating, or updating agent documentation, module docs, roadmaps, or AGENTS.md. Understands the full .augment/, agents/, and copilot-instructions structure.
日本語の概要は準備中です。原文の説明を表示しています。
Use for an adversarial red-team / blue-team / auditor review of an AI agent's CONFIG + behaviour (rules, skills, MCP, hooks, permissions) — attack-chain → defensive-gap list, not a code audit.
日本語の概要は準備中です。原文の説明を表示しています。