Use when reviewing UI for accessibility — WCAG 2.2 AA, keyboard nav, focus, ARIA, contrast, screen-reader semantics — even on 'is this a11y-OK?' or 'mach das barrierefrei'.
日本語の概要は準備中です。原文の説明を表示しています。
Use to run a depth-bounded self-correction loop (attempt → critic verdict → re-attempt) as a tunable test-time compute knob — a do-and-judge specialisation, default off, capability-gated.
インストールする前に、エージェントに与えられる指示の中身を確認できます。
verification.recursive) is not off for the
active host, or the user asks for an extra self-correction pass.Do NOT use when:
docs/benchmark.md: strong-host
discipline lift is null).verification.recursive: off for this host.Land a verified result by running a depth-bounded
attempt → critic verdict → conditional re-attempt loop, with depth as
the only compute knob, every level budgeted, and the loop default-off
until a benchmark gate authorises it per host.
Disposition (2026-07-28, honest null — TERMINAL). The recursive- verification benchmark resolved as a published null (no measured lift over single-pass verification; see
docs/benchmark.md).verification.recursivestays default-off bound to that null; scheduled for removal at the next major unless external evidence appears first. This skill is NOT sold as a quality-lift mechanism.
Scope of that null — what it does NOT close. It measured a loop that ADDS a critic to judge whether an attempt was good enough (deterministic scorer-as-critic,
max_depth=1, weak host,capH-debug). It is not evidence about retrying on a check that has already returned red: there the verdict is deterministic and in hand, and no critic is introduced, so there is nothing for the null to be null about. The measurement's own decisive finding inverts in that case — recursion was redundant because the first attempt already passed the critic 72% of the time, so cost scaled with every task while benefit sat in the ~28% tail; a red-triggered retry costs nothing on the passing majority and fires only on the tail. What the null DOES bind, for any such loop, is the falsification shape: pre-register the reduction it must deliver, and revert rather than narrate if it does not.
A CROSS-MODEL CRITIC NEVER RUNS ON THE SAME MODEL + CONTEXT AS THE ATTEMPT.
SAME-MODEL SELF-CRITIQUE IS A DISCIPLINE PASS ONLY — NEVER A CAPABILITY CLAIM.
Inherited from subagent-orchestration:
same model + same context = same blind spots. A same-model self-critique
at depth 1 is allowed only when explicitly flagged as a discipline
(not capability) pass — it can catch a skipped step or a scope-creep, but
it shares the attempt's blind spots and must never be sold as a capability
lift. Cross-model recursion (a different vendor as critic) is the
cross-vendor variant and obeys the Iron Law by construction
(critic model ≠ attempt model).
attempt₀
→ critic verdict (accept | revise: <reason>)
→ accept → done
→ revise → attempt₁ (reads attempt₀ + the verdict as context)
→ critic verdict
→ … → depthₙ
Each level reads only the prior attempt plus the critic's verdict —
never the full history — mirroring the read-your-own-output-and-decide
pattern. Depth n is the tunable compute knob, hard-capped by
verification.max_depth.
The loop stops at the first of:
accept — the critic accepts the attempt.max_depth reached — verification.max_depth (default 1).verify-budget exhausted — each re-attempt is one budgeted unit
per verify-budget; a
required-but-unrun verification is a surfaced safety gap, never a
silent pass.Stop conditions are deterministic so the loop can never run unbounded — there is no open-ended "keep trying" branch.
Configured in .agent-settings.yml; documented in
agent-settings:
| Key | Default | Effect |
|---|---|---|
verification.recursive | off | off = inert; ask = ask once before looping; on = loop silently up to max_depth. |
verification.max_depth | 1 | Hard cap on correction rounds. 1 = a single critic pass (effectively inert beyond one review) until a benchmark gate authorises more. |
Ships off. The per-host shipped default flips only on a passing
capability-axis benchmark cell (see Procedure step 4); an honest-null
keeps it off.
Resolve verification.recursive for the active host. off → no-op.
Confirm the task is non-trivial (above the verify-budget change-size
floor) and that the host plausibly has headroom — skip otherwise.
Read .agent-settings.yml (subagents.judge_model). A cross-model
critic must satisfy the Iron Law. A same-model depth-1 pass is allowed
only when explicitly flagged as a discipline pass; surface that framing.
Run attempt → verdict → conditional re-attempt, counting each
re-attempt against verify-budget, until a deterministic stop condition
fires. Under verification.recursive: on surface depth + spend in one
line; under ask, ask once before the first re-attempt.
The shipped default is set by the bench:ab gate
(orchestration-benchmark-gate,
gateVerdict / resolveShippedDefault): on/ask only on a host whose
capability-axis cell passed; off otherwise. A discipline-only lift
does not authorise a flip — that question is already answered by the
existing rules. Never flip a default without its own passing cell.
Follow the output format. Never present a recursion result as a capability gain or a frontier-model comparison.
max_depth above 1.off
there.accept / max_depth / budget / no-progress).max_depth and the stop conditions are
hard caps.| Task | Skill / context |
|---|---|
| Mode selection, the judge Iron Law | subagent-orchestration |
| Per-pass cost budgeting | verify-budget |
| Shipped-default gate mechanism | orchestration-benchmark-gate |
| Cross-vendor critic (different vendor) | ai-council |
| Completion evidence | verify-completion-evidence |
まだレビューはありません。使ってみた感想をお寄せください。
概要と使いどころ
Use when reviewing UI for accessibility — WCAG 2.2 AA, keyboard nav, focus, ARIA, contrast, screen-reader semantics — even on 'is this a11y-OK?' or 'mach das barrierefrei'.
日本語の概要は準備中です。原文の説明を表示しています。
Use when defining or auditing the activation event — aha-moment selection, retention correlation, falsifiable definition. Triggers on 'what is our aha moment', 'redefine activation'.
日本語の概要は準備中です。原文の説明を表示しています。
Use when capturing an architectural decision — file naming, next ADR number, Status / Context / Decision / Consequences, index regen; fires even without saying 'ADR'.
日本語の概要は準備中です。原文の説明を表示しています。
Adversarial critique — devil's advocate, stress-test, honest teardown ('poke holes', 'be brutal', 'was hältst du davon'); explicit request only. Routine code or design review → code-review.
日本語の概要は準備中です。原文の説明を表示しています。
Use when reading, creating, or updating agent documentation, module docs, roadmaps, or AGENTS.md. Understands the full .augment/, agents/, and copilot-instructions structure.
日本語の概要は準備中です。原文の説明を表示しています。
Use for an adversarial red-team / blue-team / auditor review of an AI agent's CONFIG + behaviour (rules, skills, MCP, hooks, permissions) — attack-chain → defensive-gap list, not a code audit.
日本語の概要は準備中です。原文の説明を表示しています。