Use when the user requests integration testing, feature validation, or test plan execution
日本語の概要は準備中です。原文の説明を表示しています。
Orchestrate work as a cyclic directed graph of subagent-executed nodes with transition criteria on edges, budgeted cycles, and per-node gates (runtime verification, metric, or artifact). Use when the user says "workgraph", "run this as a graph", "graph this work", or when a goal has parallel branches, feedback loops, or mixed gate types that a linear loop cannot express. Every node is sized to a 64k context budget, so small goals run as a two-node do/verify graph and large goals scale out without caps. For clock-driven or metric-driven single loops, timeboxed-iterating and autoresearch remain lighter.
インストール方法を見るインストールする前に、エージェントに与えられる指示の中身を確認できます。
Work expressed as a directed graph. Nodes are units of work executed by subagents. Edges carry transition criteria. Cycles are allowed — and budgeted.
You are the orchestrator: router, scheduler, gatekeeper, scribe. All productive work happens inside subagents. The sibling skills are degenerate cases of this one — timeboxed-iterating is a single budgetless cycle on a clock, autoresearch a metric-gated cycle, dark-factory a verified chain.
The one sizing rule: every node fits a 64k context budget. The planner
estimates each node from the task and sizes it to that budget. A goal that
fits one node is a two-node graph (do → verify) and runs with no further
ceremony; a very large goal is a large graph. Nodes are sized to the
budget, never to a target count.
Plan ──► lint ──► ┌─► route ─► schedule ─► dispatch ─► gate ─► log ─┐
│ │
└────────────── graph still live ─────────────────┘
└── exit signal ──► final summary
1. YOU DO NO WORK. You route, schedule, dispatch, gate, and log. Exploring
the repo, reading source, or researching the goal IS work. Nothing else.
2. THE GRAPH DIR IS GROUND TRUTH. graph.md is current state (edited in
place, rows never deleted); log.md is append-only and never rewritten.
3. EVERY CYCLE HAS A BUDGET. An unbudgeted loop is a structural defect.
4. NODES NEVER TALK TO EACH OTHER. Artifacts on edges, via the graph dir, only.
5. INTEGRATION IS SERIALIZED. One merge into the trunk at a time, gated.
Your entire tool surface: clock checks (date +%s), graph-dir reads and
writes, git worktree add/remove and clean git merge into trunk,
dispatching subagents, reading their return files, and gate spot-checks —
re-running a verification command to test a subagent's claim is GATING, not
work. "Let me just quickly check something" → no. Dispatch a subagent.
Degrees of freedom are split on purpose:
Node — one unit of work a single subagent can complete in one dispatch
inside the context budget: id, charter (one sentence), gate,
verification (commands or criteria), est (the planner's context
estimate), scope (mutable and frozen paths; empty mutable = read-only
node), heavy? (needs a GPU or an exclusive machine resource), status,
visits.
Context budget — 64k tokens per visit. The planner estimates each
node's context from the task: what it must read, what its tools will
print, and room to work. One number, made quickly, with margin. A node
that does not fit is split; a worker that returns BLOCKED for running
out of context is split by the planner before its next visit. That is
the whole feedback loop.
A node is split only because it does not fit, enables parallel work on disjoint scope, or bounds a cycle. Nothing else; a node never gets its own separate verify node, its gate does that, and the only standalone verify node is the skeleton's join.
Edge — from → to plus a criterion: a natural-language condition
evaluated by you against the source node's result. Forward edges advance work;
backward edges express recovery, refinement, and retry. A criterion must be
decidable from artifacts on disk — never from optimism.
Statuses — PENDING → READY → RUNNING → VERIFIED | FAILED, plus
ESCALATED, and SPLIT for a node replaced by expansion. One
status per node. Non-roots start PENDING; dispatch sets RUNNING.
Statuses are re-entrant: an incoming edge may send a VERIFIED or FAILED
node back to READY (regression, new evidence, refinement). visits
counts dispatches, first visit = 1; the cycle budget is a separate counter
on the cycle. History is never rewritten.
Gate types — assigned per node at plan time:
| Gate | Passes when | Lineage |
|---|---|---|
verify | Software runs; evidence per command: CHECK/COMMAND/EXPECTED/ACTUAL/RESULT; zero regressions | dark-factory |
metric | Extraction command run by you; first gated value sets Best:, afterwards strictly better than Best: | autoresearch |
artifact | Named path exists, is non-empty, git log -1 -- <path> names a commit, git status -- <path> clean | timeboxed-iterating |
Cycles — every cycle in the graph carries:
A cycle without both fails lint.
| Input | Required | Example |
|---|---|---|
| Goal | yes | "harden the importer", "optimize val_bpb", "build the TUI" |
| Repo / scope | yes | one git repo, e.g. ~/code/pace, mutable and frozen paths |
| Worker | yes (default: the harness's own subagent) | any agent that can read and write files, run a shell with git, and return text; cap its turns where the harness allows |
| Compute slots | yes (default 1) | 1 — how many heavy nodes may run concurrently |
| Node budget | optional (default 64k) | context per visit |
| Duration | optional | "overnight" = 8 hours; sets a deadline signal |
| Visit cap | optional (default 10 × nodes at plan) | hard stop on total dispatches, inherited by every added node |
| Metric spec | if any metric nodes | name, direction, extraction command |
| Focus / constraints | optional | recorded verbatim, bounds node charters |
Ask once, up front, for anything missing. After planning, never ask again.
Headless (no user present): goal or repo missing → write what is missing to
log.md and exit BLOCKED; every other input takes its default.
~/.harness/workgraph/<slug>/ — ground truth for the whole run:
graph.md — one row per node + edges + cycles (you alone write it)
nodes/ — one file per node: charter, gate, scope, verification, est
inputs/ — one file per visit: <node>-v<visit>.md, what to read and why
log.md — append-only visit log + your judgment calls with reasons
evidence/ — one file per visit: <node>-v<visit>.md, raw gate evidence
artifacts/ — node outputs passed along edges (reports, specs, diffs)
escalations/ — one file per escalated node
wt/ — git worktrees for mutating nodes: wt/<id>
mkdir -p the directory as the first action after choosing the slug;
never reuse an existing slug directory.
graph.md stays one line per node so a large graph never fills your own
context; when routing, read only the READY and RUNNING rows.
# Workgraph: <goal>
- Slug: <slug> Started: <unix ts> Deadline: <unix ts or none> Compute slots: <n>
- Trunk: <branch checked out at start>
- Budget: 64k Visit cap: 20
- Best: <metric value or none>
## Nodes
| id | status | visits | gate | heavy | est |
|----|--------|--------|------|-------|-----|
| build | READY | 0 | verify | no | 40k |
## Edges
- <from> → <to> [forward|backward]: <criterion>
## Cycles
- <id>: <node list> budget: 0/3 exit: <edge>
nodes/<id>.md holds, as headed sections: charter, gate, scope (mutable,
frozen), verification (one command per line), est, split reason.
Pass subagents paths, never contents. Each visit gets a one-screen input file; the subagent may read beyond it when needed.
Planning intelligence comes from subagents — you never explore the repo, read source, or research the goal in your own context; code knowledge arrives in a planner's return. You brief, lint, accept or redispatch, write.
graph.md (rows, edges labelled forward or backward with
predicate criteria, cycles with budgets and exit edges) and one
nodes/<id>.md per node (charter, gate, scope, verification, est,
split reason: budget | parallel | cycle). It
starts from the two-node skeleton do → verify and adds nodes only
when the budget forces a split or disjoint scope allows width. Nodes
with disjoint mutable scope and no artifact dependency are siblings
under the join, never a chain. You never transcribe the plan; you lint
it in place. Redispatch once with the lint list; a second failure
writes escalations/plan.md and exits BLOCKED.est is missing or over budget → splitdo → verify pair is
exempt)metric node with no metric spec → planner must use verify or
artifact, or emit the spec from repo docsdo → verify pair is a valid graph — run it as is.
If the goal has no parallel scope and no recovery cycle beyond that
pair, say so to the user and offer the matching sibling skill; do not
block on it.READY, fill the header (start, deadline, compute slots,
budget, visit cap, trunk). Render the initial mermaid overview (next
section) — mandatory. This step is header lines, not prose.metric subgraph needs a baseline
node run and gated first — there is no best without it.Planning is the only phase where user interaction is allowed.
graph.md is the machine ledger; the user gets a picture — a mermaid
flowchart appended to log.md under a ## mermaid entry, and shown in
chat only when a user is present. Rendering never ends your turn: in a
headless run, emit no chat text before the final summary; everything goes
to the graph dir and the loop continues.
×N suffix. One diagram, one compact legend
line, no legend walls. Fixed palette — same classes in every render:flowchart LR
spec --> build["build ×2"] --> gate
gate -- FAIL --> build
gate --> merge
docs --> merge
class spec,docs verified
class build running
class gate,merge pending
classDef pending fill:#9e9e9e,color:#fff
classDef ready fill:#1e88e5,color:#fff
classDef running fill:#fb8c00,color:#fff
classDef verified fill:#43a047,color:#fff
classDef failed fill:#e53935,color:#fff
classDef escalated fill:#8e24aa,color:#fff
Legend line: gray PENDING · blue READY · orange RUNNING · green VERIFIED · red FAILED · purple ESCALATED
First tick: skip Route and schedule every root marked READY. After every
node return, evaluate its out-edge criteria against the evidence file,
the named artifacts, and git state. Every satisfied criterion fires; a
PENDING, VERIFIED, or FAILED target → READY (never RUNNING,
ESCALATED, or SPLIT).
None satisfied → leave the status; do not invent an edge. A backward edge
increments its cycle's counter; counter over budget → the source node is
ESCALATED, see Escalation. Log every routing decision and the criterion
that fired.
Pick which READY nodes to dispatch now. No fixed width — weigh:
heavy nodes. A heavy node dispatches only when a
slot is free. Cheap nodes may overlap a heavy run freely.Before each dispatch: increment visits, create the worktree for a
mutating node if absent (wt/<id>, branch wg/<slug>/<id> from trunk;
reuse it on re-entry), set RUNNING, start the worker in the worktree
(cwd if the harness has one, else cd as the prompt's first line), and write
inputs/<node>-v<visit>.md: the charter, the acceptance criteria, scope,
incoming artifact paths, and on re-entry the check that failed last visit.
One screen. Read-only nodes run in the shared tree and must not write;
git status clean is their check.
Workers may run in the background; wait for their return however the
harness allows. Where the harness can observe and kill a worker, one
silent for 30 minutes is killed and marked FAILED; its
backward edge fires, or it escalates if there is none. The visit cap is a
hard stop: reaching it ends the run at the next Route.
One prompt per node — fill the brackets, keep the structure:
You are executing ONE node of a work graph. Attempt <N>.
Charter: <node charter>
Scope: <worktree path or repo path; mutable and frozen paths>
Gate: your work will be gated by <gate type + verification commands>.
Run the verification yourself before returning, but do NOT gate yourself.
Read first: <graph dir>/inputs/<node>-v<visit>.md. It lists what this node
needs. Read more if the work requires it.
<on re-entry: Prior visit failed at: <check>. Start there.>
Context budget: <budget> tokens. Cap every command's output (`| tail -n 200`
or equivalent); read large files by range, never whole.
Rules:
- Stay inside your charter. One node, nothing else.
- Mutating node: commit before returning; confirm with git status.
Read-only node: write nothing outside <graph dir>/artifacts/.
- Temp files to /tmp only.
- If blocked, report exactly what is blocking — do not improvise around it.
Return: write <graph dir>/artifacts/<node>-v<visit>-return.md (30 lines
max): what you did, commit hash, verification output location, and any
new work you believe this graph is missing (proposed nodes/edges, one
line each). Then print one line: DONE | BLOCKED <reason> | FAILED <reason>.
The return file is the end of a visit. Chat text, harness
notifications, or a worker that merely stops are not; gate only when
artifacts/<node>-v<visit>-return.md exists. A worker that exits without
one is FAILED at the gate. No deadline awareness for subagents. Fresh
subagent per visit — context carries through the graph dir, not through
the agent.
Per the node's gate type, on evidence you check:
git log -1, git status, and re-run at least one
verification command yourself. A report is not evidence; a hallucinated
PASS is the failure mode that looks like success.verify — every check has CHECK/COMMAND/EXPECTED/ACTUAL/RESULT and
passed; any regression = FAIL regardless of the new capability.metric — run the extraction command yourself, never read the raw
log body; empty output = FAIL. The first gated value sets Best: and
passes; afterwards strictly better than best = pass, equal, worse, or
crash = FAIL.artifact — the named path exists, is non-empty, and is committed in
the node's worktree (per the gate table).Integration of worktree branches into the trunk is the one serialized section:
graph.md. One git merge --no-ff at a
time, in the order branches gate-pass.VERIFIED node whose mutable scope intersects
the merged paths. Goal-level suite regresses →
git revert -m 1 the merge, then the merged node's backward edge
fires. Only a sibling node's own checks regress → no revert; that
sibling's backward edge fires and it re-enters on top of the new trunk.
Nothing else merges until resolved.<branch> into trunk", gate verify
with both nodes' verification, inputs = both diffs, worktree from trunk,
in a retry cycle of budget 2, then escalate) and re-gate the merged
result. You never
resolve conflicts by hand.Append one line to log.md per visit: date +%s, node, visit number,
trigger edge, verdict, one-line summary, evidence path. Judgment calls
(width changes, exit reasoning, graph edits) get their own entry: what,
why, alternative rejected.
Subagents propose missing nodes and edges in their returns; you decide.
Append accepted ones to graph.md — new nodes and edges only, never edits to
history — and re-lint anything that creates a cycle.
A split of a node that ran out of context happens here too, as new nodes
replacing the old (old node marked SPLIT, never deleted). Stall behavior
follows: a graph that goes quiet while the goal is unmet gets one
ideation node (inputs = graph.md, nodes/*.md, last 80 lines of
log.md; returns 5-10 concrete new nodes; in a cycle of budget 2) rather
than an early exit. Quiet twice after ideation is quiescence.
There is no single stop rule. Signals to weigh:
RUNNING, no criterion fires, no expansion worth
adding. The natural end.date +%s against it before every scheduling round. Let in-flight nodes
finish; dispatch nothing new past it.READY nodes no longer serve the goal.
Legitimate, but the bar is high: write the justification in log.md
before acting on it.The anti-exit discipline of the sibling skills still applies — every one of these thoughts is a trap:
| Thought | Instead |
|---|---|
| "Good enough to show the user" | Criteria still fire and time remains → schedule. |
| "The graph has mostly converged" | Mostly ≠ quiescent. Route again. |
| "Remaining nodes are too hard" | Hard is what escalation is for, not exit. |
| "Let me just fix this bit myself" | No. That is a node. Dispatch it. |
| "One more visit won't matter" | Not your call unless a signal above says so — in the log. |
Exit without a logged reason is a protocol violation.
On cycle budget exhaustion, irreconcilable merge, or structural flaw:
~/.harness/workgraph/<slug>/escalations/<node>.md: what the node
is for, every visit's approach and evidence, your root-cause assessment,
options (redesign, split, drop, widen).ESCALATED; the rest of the graph keeps running unless it
depends on the escalated node.escalations/trunk.md and never wait on an answer.After exit, append to log.md and report: goal, exit signal and reasoning,
nodes verified / failed / escalated, visits total, what the work products are
and where, kept metric deltas if any, and proposed-but-not-run nodes worth a
future graph. Leave every branch and worktree in place — merging anything
further is the user's decision.
log.md doesn't say whygraph.md to route when only READY/RUNNING rows matterまだレビューはありません。使ってみた感想をお寄せください。
概要と使いどころ
Use when the user requests integration testing, feature validation, or test plan execution
日本語の概要は準備中です。原文の説明を表示しています。
Use when the user wants to systematically fix AI code slop — duplicated logic, over-engineering, silent error swallowing, convention drift, cargo-cult patterns, and other LLM-introduced architectural decay — over a specified duration
日本語の概要は準備中です。原文の説明を表示しています。
Produce a researched long-form article from a topic prompt via an orchestrated pipeline - research agent (first-person sources, working-definition gate), narrative-architecture outline, writer/cold-reviewer loop with an explicit ACCEPT/REVISE verdict contract, then a catalog-deslop pass with a regression gate. The orchestrator dispatches subagents only; the writer never judges its own draft. Use when the user says "article factory", "write an article about X", "run the article pipeline", or asks for a researched long-form piece produced end-to-end. For essays and micro posts in the user's own voice without a research stage, use the prose skill instead.
日本語の概要は準備中です。原文の説明を表示しています。
Runs autonomous keep/discard experiments on a codebase to optimize a single metric for a fixed duration, in the style of karpathy/autoresearch. Use when the user says "autoresearch" (optionally with a focus, e.g. "autoresearch the optimizer"), asks to run experiments on a repo overnight, to hill-climb or optimize a metric autonomously, or points at a repo with a karpathy-style program.md.
日本語の概要は準備中です。原文の説明を表示しています。
Create custom modules for [Harbor Boost](https://github.com/av/harbor/tree/main/boost), an optimizing LLM proxy. Use when building Python modules that intercept/transform LLM chat completions—reasoning chains, prompt injection, structured outputs, artifacts, or custom workflows. Triggers on requests to create Boost modules, extend LLM behavior via proxy, or implement chat completion middleware.
日本語の概要は準備中です。原文の説明を表示しています。
Systematically explore and test any software project (CLI, API, Backend, Library, etc.) to find bugs, usability issues, and edge cases. Produces a structured report with full reproduction evidence (exact commands, inputs, logs, and tracebacks) for every issue.
日本語の概要は準備中です。原文の説明を表示しています。