Use when the user requests integration testing, feature validation, or test plan execution
日本語の概要は準備中です。原文の説明を表示しています。
Run a task iteratively over a user-specified duration by dispatching subagents. The orchestrator stays strictly linear and time-checked, but each iteration fans out MULTIPLE subagents in parallel over independent units (and those subagents may fan out further). Shared state — a compact progress digest and a growing environment cheatsheet — lives on disk and is passed to every subagent BY REFERENCE, killing the per-subagent rediscovery tax. Every subagent first reads an initialiser preamble that points it at that shared state and makes it write findings back, so the cheatsheet populates itself; role prompts are produced ONCE on the filesystem by a single scaffold command and dispatched by path, never re-typed per dispatch. Subagents persist their results to disk and return only a tiny status, so the orchestrator's context stays small and lasts. The orchestrator classifies the goal as FINITE or OPEN-ENDED: finite lists may finish early; open-ended goals use the full duration, never idle, never manufacture busywork. Use when the user gives a task and a duration and wants it ground out iteratively by subagents over that time.
インストール方法を見るインストールする前に、エージェントに与えられる指示の中身を確認できます。
Take a goal and a duration and grind it out with subagents. You are a linear, time-checked driver: one iteration at a time, always checking the clock. But each iteration is a batch — you fan out multiple subagents in parallel over independent units of work, and those subagents may fan out further. You never author anything yourself. Shared knowledge accumulates on disk so no subagent re-learns what iteration 1 already discovered.
You are the orchestrator. You do exactly five things:
date +%s before every batch; the deadline is the brake.done transition after the measurement gate.You do no authorship and no analysis: no code edits, no DB queries, no PR creation, no HTML/report generation, no "quick fixes," no reading the codebase to figure something out. Every productive act happens inside a subagent — even under context pressure, even after a blowout (this is the Orchestrator-Only Line, and it is the fix for the failure where orchestrators burned their scarcest resource doing delegable work). Your context is reserved for the dispatch loop. If you catch yourself doing anything other than checking time, managing the on-disk files, planning, and dispatching — stop. That work belongs in a subagent.
Pure dispatch — never ingest large content (E3). Orchestrator context is the
scarcest resource on a long run; exhausting it forces premature restarts. So you
never read a full artifact, a full subagent transcript, or a large return into
your own context. Subagents persist their results to disk — their unit file, the
digest, the cheatsheet — and return only a TINY structured status (e.g.
unit X: done, commit <hash> or measure X: PASS — <one line>). You read only the
compact digest tail plus those one-line statuses. Prefer having subagents
append their own digest/unit-file entries so you are not even reproducing that
content. Per-iteration footprint stays minimal: clock check → glance at the digest
tail → dispatch by reference → record/confirm a one-line status. If a return is
about to dump an artifact or a wall of prose into your context, that is a Red Flag
— it belongs in the unit file, fetched by whoever needs it, not by you.
1. THE CLOCK (OR THE USER) DECIDES WHEN AN OPEN-ENDED RUN STOPS. Not you, not a subagent, not a "feels done."
2. A FINITE LIST DECIDES WHEN A FINITE RUN STOPS. Exhaust the list → stop. Finishing early is CORRECT, not failure.
3. THE ORCHESTRATOR NEVER AUTHORS. It checks time, manages disk state, plans, dispatches, reads returns. Nothing else.
4. EACH ITERATION FANS OUT; INDEPENDENT WORK RUNS IN PARALLEL. Serial dispatch of independent units is the dominant, forbidden loss.
Law 1 preserves this skill's one genuine strength — an open-ended run must not stop early just because progress "looks good." Law 2 narrows that strength: it does not apply to a finite, exhaustive spec or a finite item list, where padding the clock after the list is done is pure waste. You must classify the goal (Phase 0) to know which law governs.
Violate any of these and you are running a slow serial loop, not this skill:
<harness>/prompts/ once and dispatch it by path too. Every dispatch is a
role-prompt PATH plus a short unit delta — never a re-typed template. (E1/E4/D3/C2)If the duration is vague, interpret "overnight" as 8 hours. If the goal is genuinely ambiguous, ask one question. Otherwise pick and proceed.
digraph timeboxed {
rankdir=TB;
node [shape=box];
classify [label="Phase 0: Classify goal\nFINITE vs OPEN-ENDED\nrun scaffold: workspace + prompts/ + digest + cheatsheet + clock" shape=doublecircle];
front [label="Phase 1: Setup + Front-load\nscaffold produced prompts/ ONCE;\nsuccess-bar pilot (numeric OR textual) + enabling infra +\nscope/validity bar + repo sync\n(seeds cheatsheet; then it grows dynamically)"];
brake [label="Brake + Stop-guard\n(deadline? finite list done?\nremaining time fit a useful unit?)" shape=diamond];
plan [label="Phase 2a: Plan iteration\nread digest; select independent units;\npartition parallel/serial; size a model each;\ncap batch at concurrency limit"];
dispatch [label="Phase 2b: Dispatch BATCH in parallel\neach subagent: role-prompt PATH + unit-file path +\ngap line (short delta) — NO inline templates"];
collect [label="Phase 2c: Read TINY statuses (never full artifacts)\nspot-check each commit via git log"];
consol [label="Phase 2d: Consolidate (minimal footprint)\nsubagents already appended to digest+cheatsheet;\nflip finite-list done only after the success gate PASSes"];
done [label="Phase 3: Stop\nfinal summary" shape=doublecircle];
classify -> front -> brake;
brake -> plan [label="OPEN-ENDED and a useful unit fits\nOR FINITE and list not exhausted"];
brake -> done [label="deadline passed / user stop /\nFINITE list exhausted /\nno useful unit fits remaining time"];
plan -> dispatch -> collect -> consol -> brake;
}
Every path returns to the brake check. done is reachable only from the
brake check — there is no quality-based or "feels complete" exit anywhere in the
graph. Where the diagram and the prose disagree, the Stop Machine below is
authoritative.
This is the first decision and it governs the entire stop behavior. Write the classification into the digest header; it is not revisable on a whim.
If a goal is finite in one dimension and open-ended in another (e.g. "review these 10 files for bugs" — finite file list, open-ended depth per file), treat the enumerable dimension as the list and the depth as open-ended within each item: the run ends when the list is covered to a defined depth or the clock fires, whichever first.
You do not hand-build the workspace or paste template text. ONE command produces
the entire ready-to-run workspace at ~/.harness/timeboxed/<slug>/ from a few
variables — see The Scaffold Command below for the full contract. Run it now:
bash <skill-dir>/scaffold/init.sh \
--slug <slug> --goal "<goal>" --mode <finite|open-ended> [--duration <e.g. 4h|90m>]
It creates and variable-substitutes:
progress.md — the STATE DIGEST: compact, single source of truth for
volatile values (iteration count, per-unit status, commits,
the finite item list + what's done). Subagents append one-line
statuses; you read only its tail. NOT a pile of ledgers.
run-card.md — the static RUN CARD: run identity + workspace map + resume
pointer. Holds NO volatile values (those live only in
progress.md). Read it once to orient.
cheatsheet.md — persistent, growing: environment recipes, working commands,
tool quirks, known gotchas, auth workarounds. Seeded empty of
findings; read by every subagent, appended to by every subagent.
prompts/ — role prompt files, produced ONCE by the scaffold from the
filesystem templates in `scaffold/`: initialiser.md (read FIRST
by every subagent), builder.md, sub-subagent.md, measurement.md.
Dispatches pass these BY PATH, never re-typed.
units/<unit-id>.md — one file PER UNIT: its scope on dispatch, its full result on
return. Per-unit files exist so parallel subagents never write
to the same file (no write contention), and so full results live
on disk instead of in the orchestrator's context.
After scaffolding, fill the two placeholders the scaffold left in progress.md:
the concurrency cap and the real success bar (numeric OR textual — Phase 1
pilots it; write none if the goal has no statable bar). Dispatch a subagent for
anything non-trivial — you author nothing.
Create the target repo context: the artifact is always under git (existing repo
in place, or mkdir ~/code/<slug> && git init for from-scratch work). Every
subagent commits; your spot-checks depend on it.
date +%s, compute the deadline epoch, write both to the digest. You check it at
the top of every iteration (the brake), before dispatching the batch — never
after, never "when convenient."
Do this before the main loop, in the first iteration(s). You author nothing — the front-loading work itself runs via dispatched subagents (Iron Law 3).
The scaffold command (Phase 0b) already produced the four role prompts under
<harness>/prompts/, variable-substituted:
initialiser.md — the preamble EVERY dispatched subagent reads first.builder.md — the builder role prompt.sub-subagent.md — the fan-out role prompt.measurement.md — the independent measurement role prompt.You do not re-type or re-compose them. Thereafter every dispatch hands the subagent a role-prompt PATH plus a short unit delta (2b). Re-typing one burns your context (the resource this skill protects) and is a Red Flag.
The general rule (E1): write any reusable prompt to disk ONCE, then reference it
by path in every identical situation. The scaffold covers the four standard
roles. If you ever find yourself composing a specialized prompt for a recurring
unit type, do not paste it into each dispatch — write it into <harness>/prompts/
one time and dispatch it by path exactly like the standard roles.
Front-load only what is a real ENABLER and provably exists — not a speculative "run the whole task once to see what breaks." A pilot may surface no issues, and issues surface continuously, not at t=0. In priority order:
none and
skip the gate — but prefer to state a bar wherever one exists.cheatsheet.md; the rest grows dynamically thereafter.Everything else is dynamic, not front-loaded. Environment recipes beyond the
initial standup, tool quirks, gotchas, and reusable findings are discovered
continuously by subagents that read prompts/initialiser.md first and append what
they learn to the cheatsheet (D1). Do not try to enumerate every gotcha up front
from one test run — that is what the initialiser+cheatsheet loop is for. Keep
B8's real point (never defer a genuine enabler to the final hour) without turning
issue-discovery into a t=0 test run.
Record in the digest that front-loading ran and what it established. The cheatsheet is seeded; from here it grows itself, batch over batch.
The loop is linear: one batch at a time, brake-checked at the top. Within a batch, subagents run in parallel.
Run date +%s, compare to the deadline, and evaluate the Stop Machine. If
any stop condition holds → Phase 3. Otherwise continue to 2a.
A batch already in flight when the deadline passes soft-stops: let its subagents finish and consolidate their returns, then stop. Never strand committed work unrecorded.
Read the digest (the compact one, not a heap of ledgers — B3). From it, select the next set of units — the smallest pieces of real work that move the goal. For a FINITE run, units are the next unclaimed items from the list. For an OPEN-ENDED run, units are the next highest-leverage angles.
Partition units into parallel vs. serial (C1 / B1) — the independence rule:
Respect the concurrency cap. Never put more than the platform's cap of subagents in one batch. If you have 20 independent units and a cap of 6, dispatch 6, let them return, dispatch the next 6. Over-cap dispatches fail silently and you'll redispatch blind. The cap is in the digest header.
Size a model per unit (C3). For each unit choose the smallest model that can do it well:
| Unit character | Model |
|---|---|
| Mechanical / narrow (rename, mechanical edit, run a known command, format, apply a known fix, scrape one page) | Small / cheap model |
| Standard implementation, moderate reasoning | Mid model |
| Deep design, novel analysis, hard debugging, judgment | Orchestrator-tier |
Hard cap: never dispatch a model more powerful than the orchestrator itself. If you are running on a mid-tier model, "orchestrator-tier" IS your ceiling — you cannot summon a stronger model than yourself; asking for one silently fails or wastes the dispatch. Size down freely, never up past your own tier.
Write each unit's scope into its own units/<unit-id>.md before dispatch.
Dispatch all units in the batch in a single step (one message, multiple subagent calls) so they run concurrently. Every dispatch is a SHORT message — a role-prompt PATH plus a small unit-specific delta, never a re-typed template:
<harness>/prompts/builder.md (or measurement.md for
the measurement gate). That file already tells the subagent to read
<harness>/prompts/initialiser.md first, which points it at the digest and
cheatsheet by reference and makes it write findings back.<harness>/units/<unit-id>.md
(you wrote the scope there in 2a), and one line naming this unit's gap/target.That is the whole dispatch. You never paste a template, the digest, or the cheatsheet into the message (C2/D3) — you pass paths. Re-typing a template into a dispatch is a Red Flag: it burns your context, the exact resource this skill exists to protect.
The role prompts already instruct each subagent to persist its full result to disk (its unit file), append its findings to the cheatsheet, append one status line to the digest, and return only a tiny status — not a wall of prose (E3). You depend on that: your context stays small only because results land on disk, not in returns.
Do not give subagents the deadline or any time awareness. They do one unit and return. Time is your concern alone.
When the batch returns, read each subagent's tiny status line — not its
artifact, not its transcript. The full result is in the unit file if anyone ever
needs it; you do not pull it into your context. Spot-check each claimed commit
with git log (the commit is there / it is not) — a cheap on-disk check, not a
read of the produced content. Subagents hallucinate deliveries; 30 seconds of
git log saves a wasted iteration. A unit that claims a commit not in git log
is re-dispatched, not recorded as done. If a subagent tries to hand you the
artifact itself, ignore the body and read the digest tail instead.
Keep your own footprint minimal (E3): the subagents already appended their status lines to the digest and their findings to the cheatsheet. You are confirming, not re-authoring.
done transition in the digest: after the success gate passes, mark
the finite-list item done and its commit — a one-cell edit, from the tiny
statuses you were handed, never by reading the artifacts. The digest is the ONLY
place volatile values live. Do not copy counts/hashes into other files (or into
run-card.md, the static run card) — that O(n) re-sync is exactly the bookkeeping
churn this skill forbids.cheatsheet.md; parallel appends of a few lines rarely collide, but if two
batches raced and an entry is malformed, fix it in one edit). New environment
facts must survive to the next batch — that is what kills the rediscovery tax
(B2). Same for the one-line digest status appends.done in the digest ONLY after it has passed that real bar,
scoped to the new unit(s) — never a cheap pass-shaped proxy. A unit that has not
yet passed stays doing, or is re-dispatched; it is never marked done on a
proxy. The bar is applied by a separate, freshly dispatched measurement
subagent that did not produce the unit — never the producer's self-report
(dispatch it by passing the path <harness>/prompts/measurement.md plus the unit
delta — never re-type the template). The measurement subagent writes its full
evidence into the unit file and returns only PASS|FAIL — <one line>; you gate on
that one line, not on its evidence body. If the goal has no statable success
bar (pure qualitative churn, nothing identified in Phase 1), this gate does not
apply — ordinary scoped verification below is sufficient.Then return to the brake check. Do not write a between-iteration status message to the user (that is orchestrator authorship and a Red Flag). The digest is the live status page.
The run ends when, and only when, at the brake check one of these holds:
done
and no valid unit remains. Stopping here is correct (Iron Law 2) — do not
invent work to fill the remaining clock.No other stop exists. In particular these are FORBIDDEN, not stops:
sleep 1000 is a firing offense).For an OPEN-ENDED run with time left and no obvious unit: that is a stall, not a stop — go to Stall Recovery. For a FINITE run with the list exhausted: that is a stop (condition 3), immediately.
The role prompts and the workspace are NOT embedded in this skill and are NOT
hand-typed. They live as template files in this skill's scaffold/ directory, and
ONE command materialises them into a ready-to-run workspace with the run's
variables substituted in. This is the concrete form of the write-once/reference-
by-path rule (E1): the templates are authored once in scaffold/, produced once
per run on disk, and thereafter dispatched by path — never re-typed.
bash <skill-dir>/scaffold/init.sh \
--slug <slug> \
--goal "<goal>" \
--mode <finite|open-ended> \
[--duration <e.g. 4h | 90m | 2h30m>] \
[--force]
--slug — workspace name; the workspace is ~/.harness/timeboxed/<slug>/.--goal — the one-line goal (substituted into every prompt + the digest header).--mode — finite or open-ended (Phase 0 classification; drives the stop law).--duration — the timebox when timed (4h, 90m, 2h30m, or a bare number =
hours); the script computes and records the deadline. Omit for an unbounded run.--force — overwrite an existing non-empty workspace. Without it, the command
refuses to clobber one (safe to re-run).It prints every path it created. After it runs, fill the two placeholders it left
in progress.md: the concurrency cap and the real success bar (numeric OR
textual, or none).
prompts/initialiser.md, prompts/builder.md, prompts/sub-subagent.md,
prompts/measurement.md — the four role prompts, {{HARNESS}}/{{GOAL}}/…
substituted. Dispatched BY PATH; never re-typed.progress.md — the digest (single source of truth; header + finite list +
iteration log).run-card.md — the static run card (identity + workspace map + resume pointer).cheatsheet.md — seeded with section headings, empty of findings; self-populates.units/ — one file per unit at dispatch time.scaffold/)unit <id>: done, commit <hash>.measure <id>: PASS|FAIL — <one line>.To change a role's wording, edit the template in scaffold/ — never re-type it
into a dispatch. Structure to preserve: every role reads the initialiser first;
builders persist to disk + return a tiny status; the measurement role judges
against the real bar and returns PASS/FAIL + one line, with no artifact and no
commit.
Every thought on the left will occur to you. Do the right column instead.
| Thought you're having | What you must do instead |
|---|---|
| "These units are independent, I'll just do them one at a time" | No. Independent units go in ONE parallel batch. Serial dispatch of independent work is the dominant loss. (B1) |
| "I'll paste the progress so far into the prompt" | Pass the digest PATH. Subagents read it themselves. (C2) |
| "I'll write the builder/measurement template into each dispatch" | The scaffold wrote prompts/ ONCE; every dispatch passes the role-prompt PATH + a short unit delta. Re-typing a template burns your context. (E1/D3) |
| "I'll compose this specialized prompt fresh each time I need it" | Write it to <harness>/prompts/ ONCE, then dispatch it by path in every identical situation. (E1) |
| "I'll read the unit's artifact / the full return to see what it did" | No. Read the tiny status + the digest tail. Full results live in the unit file; ingesting them burns your context. (E3) |
| "This subagent should just return its whole write-up to me" | No. Subagents persist to disk and return a one-line status; that is what makes your context last. (E3) |
| "I'll hand-build the workspace / paste in a starter digest" | Run scaffold/init.sh — one command builds prompts/ + digest + run card + cheatsheet + units/. (E4) |
| "The success bar is a number, so text-only goals can't be gated" | The bar is numeric OR textual (rubric, acceptance description, judge verdict). The independent gate applies either way. (E2) |
| "I'll front-load every gotcha from a test run up front" | Front-load only genuine enablers + the success-bar pilot. Recipes/quirks/gotchas populate the cheatsheet DYNAMICALLY, via the initialiser every subagent reads first. (D1) |
| "The subagent can just figure out the environment" | Point it at the cheatsheet FIRST (the initialiser does this); it must not re-pay the discovery tax. (B2) |
| "Let me re-read all the ledger files to be safe" | Read the compact digest only. No re-reading the whole corpus. (B3) |
| "Let me sync the count/hash into these other files too" | One source of truth — the digest. Never fan volatile values out. (B4) |
| "I'll mass-produce now and measure at the end" | Pilot the real measurement FIRST; it gates the loop. (B5) |
| "The pilot already proved the measurement works, this unit can pass on a quick proxy" | No. Every unit gated by a real measurement must PASS IT — not a proxy — before being marked done. (B5) |
| "The subagent that built the unit can also report it passed the real measurement" | No. Dispatch a SEPARATE measurement subagent; the producer's self-report never gates done. (B5) |
| "The list is done but there's time left — find more busywork" | If FINITE and exhausted, STOP. Finishing early is correct. (Iron Law 2 / B6) |
| "Almost out of time, let me dispatch one more anyway" | If no useful unit fits the remaining time, STOP. No deadline-edge dispatch. (B6) |
| "I'll just sleep out the rest of the clock" | Never. Idle sleeping is a firing offense. (B6) |
| "I'll quickly run this DB query / make this PR myself" | No. Dispatch a subagent. You author nothing — even under context pressure. (B7) |
| "Context is tight, I'll just do it inline this once" | Especially then — inline work burns your scarcest resource. Dispatch. (B7) |
| "Infra/scope can wait till later" | Front-load it in the first iteration or it becomes a rework phase. (B8) |
| "I'll dispatch all 20 units at once" | Cap the batch at the concurrency limit or they fail silently. |
| "This open-ended run looks good enough" | Not your call. Check the clock; if time remains, dispatch. (Iron Law 1) |
| "Let me write the user a progress update" | The digest is the status page. Only the final summary goes to the user. |
| "I'll grab a bigger model for this hard unit" | Never above your own tier. Orchestrator-tier is the ceiling. (C3) |
If you catch yourself forming any opinion about whether the work is "done" on an OPEN-ENDED run — check the clock and dispatch instead.
A batch returning "nothing meaningful to do" on an open-ended goal is a stall, not a stop:
(For a FINITE run, "nothing left" is not a stall — it is the exhausted-list stop. Do not reframe a finite goal into busywork.)
prompts/… path plus a short unit delta.prompts/ once
and dispatching it by path.prompts/ was not produced by the scaffold at setup, so dispatches carry full
templates (run scaffold/init.sh).run-card.md).done without passing the real success bar (numeric OR
textual) — only a cheap proxy check ran.done on the producer's own claim, with no independent
measurement subagent dispatched.sleep to pass time.date +%s since the last batch returned.git log.If ~/.harness/timeboxed/<slug>/ already exists, reconcile disk against memory —
trust disk:
progress.md. Note the classification (FINITE/OPEN-ENDED), the deadline,
the concurrency cap, and the finite item list if any.git log in the target repo to
the commits recorded in the digest. A timeboxed(<unit-id>): … commit not
recorded in the digest means the run died between the commit and the digest
write — record it now (that's why the prefix is mandatory).done marks against actual commits. An item marked doing with a matching
commit is really done; an item marked doing with no commit is unclaimed —
re-dispatch it.prompts/. Confirm prompts/ holds the four role files
(initialiser, builder, sub-subagent, measurement). If it is missing or partial,
re-run the scaffold command (scaffold/init.sh, with --force if the workspace
exists) before dispatching — never resume by re-typing templates into dispatches.| Item | Value |
|---|---|
| Harness dir | ~/.harness/timeboxed/<slug>/ |
| Scaffold | scaffold/init.sh --slug --goal --mode <finite|open-ended> [--duration] — ONE command builds the whole workspace from filesystem templates; prints every path; safe to re-run (--force to overwrite) |
| Classification | FINITE (finite list = stop authority) vs OPEN-ENDED (clock = stop authority) — decided in Phase 0 |
| Digest | progress.md — compact single source of truth for volatile values; the live status page; orchestrator reads only its TAIL |
| Run card | run-card.md — static run identity + workspace map + resume pointer; holds NO volatile values |
| Cheatsheet | cheatsheet.md — populates itself: initialiser makes every subagent read it first + append findings back; grows every batch |
| Prompts | prompts/ — initialiser + 3 role prompts, written ONCE by the scaffold; dispatched BY PATH, never re-typed; write any reusable specialized prompt to disk once too (E1) |
| Initialiser | prompts/initialiser.md — EVERY subagent reads it first: read shared state, persist result to disk, append findings + a one-line status, return tiny |
| Unit files | units/<unit-id>.md — one per unit; parallel subagents never share a write target; full results live here, not in the orchestrator's context |
| Time check | date +%s vs deadline, at the top of EVERY iteration, before the batch |
| Iteration | A BATCH: independent units dispatched in parallel (≤ concurrency cap), each sized to a model ≤ your own tier |
| Dispatch | A SHORT message: role-prompt PATH + unit-file path + one gap line. Never a re-typed template, digest, or cheatsheet |
| Return | TINY status only (unit X: done, commit <hash> / measure X: PASS|FAIL — one line); full result is on disk. Orchestrator never ingests artifacts/transcripts (E3) |
| Independence rule | Separate files → parallel; shared file → serial round-robin; within a unit, subagents may fan out further |
| Front-loading | Scaffold prompts/ + success-bar pilot + genuine enablers (infra, scope/validity bar IF any, repo sync) in the FIRST iteration(s); recipes/quirks discovered dynamically thereafter |
| Success gate | Per unit, ENFORCED: a unit gated by a real bar (NUMERIC or TEXTUAL/qualitative) is marked done only after a SEPARATE, independent measurement subagent (never the producer) passes the REAL bar, never a proxy or self-report; no statable bar → gate doesn't apply |
| Orchestrator | Pure dispatch: checks time, manages digest+cheatsheet+prompts, plans, dispatches by path, reads tiny statuses — authors NOTHING, ingests no large content, ever |
| Model sizing | Smallest model that does the unit well; hard cap at orchestrator's own tier |
| Stop (legal only) | Deadline / user / FINITE list exhausted / no useful unit fits remaining time |
| Never | Idle sleep, manufactured bookkeeping, deadline-edge dispatch, inline authorship, ingesting a full artifact/return, over-cap batch, up-tier model, re-typing a template into a dispatch |
| Soft stop | Deadline mid-batch → in-flight batch finishes and consolidates, then stop |
まだレビューはありません。使ってみた感想をお寄せください。
概要と使いどころ
Use when the user requests integration testing, feature validation, or test plan execution
日本語の概要は準備中です。原文の説明を表示しています。
Use when the user wants to systematically fix AI code slop — duplicated logic, over-engineering, silent error swallowing, convention drift, cargo-cult patterns, and other LLM-introduced architectural decay — over a specified duration
日本語の概要は準備中です。原文の説明を表示しています。
Produce a researched long-form article from a topic prompt via an orchestrated pipeline - research agent (first-person sources, working-definition gate), narrative-architecture outline, writer/cold-reviewer loop with an explicit ACCEPT/REVISE verdict contract, then a catalog-deslop pass with a regression gate. The orchestrator dispatches subagents only; the writer never judges its own draft. Use when the user says "article factory", "write an article about X", "run the article pipeline", or asks for a researched long-form piece produced end-to-end. For essays and micro posts in the user's own voice without a research stage, use the prose skill instead.
日本語の概要は準備中です。原文の説明を表示しています。
Runs autonomous keep/discard experiments on a codebase to optimize a single metric for a fixed duration, in the style of karpathy/autoresearch. Use when the user says "autoresearch" (optionally with a focus, e.g. "autoresearch the optimizer"), asks to run experiments on a repo overnight, to hill-climb or optimize a metric autonomously, or points at a repo with a karpathy-style program.md.
日本語の概要は準備中です。原文の説明を表示しています。
Create custom modules for [Harbor Boost](https://github.com/av/harbor/tree/main/boost), an optimizing LLM proxy. Use when building Python modules that intercept/transform LLM chat completions—reasoning chains, prompt injection, structured outputs, artifacts, or custom workflows. Triggers on requests to create Boost modules, extend LLM behavior via proxy, or implement chat completion middleware.
日本語の概要は準備中です。原文の説明を表示しています。
Systematically explore and test any software project (CLI, API, Backend, Library, etc.) to find bugs, usability issues, and edge cases. Produces a structured report with full reproduction evidence (exact commands, inputs, logs, and tracebacks) for every issue.
日本語の概要は準備中です。原文の説明を表示しています。