本文へ移動
cccskills
無料GitHub で公開

system-path-simulation

Test an uncertain Anabasis path at the smallest real layer that can decide it: deterministic joins, live authoring and admission, semantic host probes, or controlled model comparisons. Use full-run rehearsals only when the question needs them; not for ordinary run review or a path already proved by its owning test.

インストール方法を見る

含まれるファイル(39)

  • SKILL.md34.7 KB
  • agents/openai.yaml391 B
  • cases/authoring-comparison.md10.7 KB
  • cases/brittle-angles.md2.6 KB
  • cases/check-one-fact.md1.7 KB
  • cases/fullrun-conditions.md5.6 KB
  • cases/layer-walk.md5.9 KB
  • cases/live-segment.md15.9 KB
  • cases/past-run-replay.md4.7 KB
  • cases/real-run-watch.md2.8 KB
  • cases/review-replay.md6.5 KB
  • cases/scenario-stub.md6.4 KB
  • cases/seeded-condition.md6.6 KB
  • cases/seeded-project.md5.9 KB
  • cases/stacked-prefix-groups.md7.6 KB
  • references/e2e-examples.md26.0 KB
  • scripts/condition-evidence.mts11.8 KB
  • scripts/credential-use.mts4.1 KB
  • scripts/difficulty-watch.mts9.3 KB
  • scripts/file-head.mts707 B
  • scripts/host-panel.mts16.2 KB
  • scripts/judge-replay.mts10.9 KB
  • scripts/pick-run.mts10.1 KB
  • scripts/position-packet.mts6.4 KB
  • scripts/prediction-note.mts1.5 KB
  • scripts/predictions.mts10.8 KB
  • scripts/process-census.mts4.0 KB
  • scripts/review-settle.mts10.4 KB
  • scripts/run-condition.mts31.4 KB
  • scripts/run-segment.mts34.3 KB
  • scripts/seed-campaign.mts38.3 KB
  • scripts/seed-kickoff.mts4.6 KB
  • scripts/segment-actor.mts2.2 KB
  • scripts/segment-loop.mts7.3 KB
  • scripts/session-env.mts1.8 KB
  • scripts/show-prompt-surfaces.mts3.7 KB
  • scripts/test-support.ts907 B
  • scripts/tool-tree.mts11.8 KB
  • scripts/workspace-changes.mts8.0 KB

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

System Path Simulation

Use this when you need to know whether a change will work in the real system, a path has not been walked by a real actor, or a possible blocker remains unresolved. This is not tied to launch or to failure. Take the smallest real slice around the question: start where the change takes effect when that directly answers it, or one or two production transitions before it when the joins are part of what must work. A fresh unused project on a familiar unchanged path does not need it.

Revamp log

This skill is being rebuilt from observed simulations, one dated entry per run. Read the entries before choosing a case. When a few simulated runs stand behind them, overhaul the skill around what they show and delete the settled entries.

  • 2026-09-15, gate overhaul (PR #683). A seeded round starts from a copied campaign that has already passed the fresh-project path, so it cannot measure what the overhaul changed: time to the first preview, the rows each refusal shows, adoption. There the real run is the condition; cases/real-run-watch.md is the new case. A steward subagent per condition shares the account's session limit: one 429 (session limit, resets 14:20) ended the audit subagent and the run's Builder in the same minute; a fresh token on the same account died the same way at 13:49. The launcher now runs one minimal Builder-slot turn before the gate and refuses with the provider's reset clause (main 69ed55b5b). Stewards went; the launching session ran its own conditions. Seeded rounds default the Built slot to scripted, so the Builder → Built Harness handover stays unexercised until a real battery; when the question is the measure stage, run --built live.
  • 2026-09-29, stewards return on an account of their own. The 2026-09-15 death was one account carrying the steward's condition, the run's Builder and the parent at once. The operator's numbered accounts in .accounts/ separate them: a steward's condition spends an account no live run and no parent is spending, so its 429 ends only itself. The rule is below.
  • 2026-09-30, three truss and esp32 stewards spent their first minutes on hand steps: a seed refused 3,466 tool-tree references, and a second five-minute --relocate run left them naming the main checkout's tree; Built pins unlike the record turned a rebuild into a measure; CLAUDE* variables were unset by hand; a transcript copy was lost to the CLI's close stub. Helpers own each step now.
  • 2026-10-01, the limit-line arm of PR #122 from esp32 campaign -20 at its partial battery i14. Two seeds and two captures, no model turn. Neither first prompt carried the line: the selector measures a republished product before it builds, so with --built no-solve the newest row is a refused claim (DISCRIMINATION_ACCEPT_REJECTED, 19 invalid control receipts) and a line that speaks for the latest battery alone cannot fire; and the readout set every carried battery aside, since products recorded before experiment-authoring/v2 (2026-09-29) are not resumable, by design. A condition on how a Builder answers a battery needs a position recorded after that cut whose seed measure is reproduced (a scripted solver over the recorded artifacts) or solved live; every esp32 campaign after the cut is an Opus full pass, so the arm's trigger has not yet existed for Opus 5.5 there. pick-run.mts should say the cut. The same evening, three live segments from that position (run-segment.mts, Opus 5.5 medium, 12 turns, 30-minute turn wall, the readout rendered from the recorded counts and labelled authored): the arm's limit line, the control without it, and B′ with i14 authored as a full pass. The arm and the control did the same thing: kept the failed task byte for byte, named its stand-in check wrong and widened it (probe.ts), and added 14 tasks to reach the 25 the task-count line asks; the control read the fail from the families line (esp32-oled 0/1) and the traces. B′ removed all 11 tasks and authored 25 new ones over a new host stand-in. So at this position the line changes nothing: the Builder's branch follows the readout's reading of the record, and the arm adds no second reading. The seed's cost: run-segment.mts copies no .toolchain, so each Builder rebuilt a 10 GB arduino tree and lost its first 30-minute turn to the wall, two of three a second; a segment whose actor compiles needs a cloned tree or a wall above the install.

Standing triggers

These change classes warrant a simulation; each earned its place in a recorded session (2026-09-01: four stewards, two changed a conclusion).

The change looks likeSteward and question
An interface where the tests use doubles — backend signal contracts, vendor/pi-claude-bridge, transportsOne layer-walk over the real sources: does each field/signal the change reads actually arrive there? Both recorded production defects (the one-hour settle fallback and Claude's MCP-prefixed tool names; earlier, the acceptance-close false warranty) were green-tested wrong joins of this class.
A removed default, backstop or widened guardOne steward asking "what fires now that this doesn't?" Removing the six-hour session cap uncovered a hidden one-hour DEFAULT_TURN_SETTLE_MS that fired two hours before the new wall could start.
A fix that claims to repair a specific past run, before the next paid runOne past-run-replay over that run's actual recorded bytes (claims, packets, terminals), plus the nearest hostile mutation in a scratch copy. Minutes of replay against run36/run39/run40 bytes beat another dead run discovering the miss live.
A recorded position whose next controller round is the questionOne seeded-condition: seed with seed-campaign.mts, capture the first prompt, run one round through run-condition.mts with the slot under test live, probe with host-panel.mts. The 10 September Astra condition answered handover, admission and scope binding in one 354-second native turn; its helper is now the runner.
A changed Judge or Epoch Reviewer prompt, schema or orientationOne review-replay: the live Judge over recorded cases with judge-replay.mts, then the live reviewer over the resulting contested rows with review-settle.mts. On 2026-09-15 the first replay refuted the chosen position in six minutes (both recorded disputes were the Judge's arithmetic) and the second settled a genuine veto against the harness.
An unlanded stack before a paid runOne triage, no model: list every decision change in the stack with no live exercise, give each one condition, and order them so the cheap condition can cancel the dear one — one fact, then layer walk, then deterministic replay, then the eight-task one-iteration rehearsal, then a live segment. On 2026-09-01 two of four listed conditions settled without a model: a writer/reader join (below) and a revert proved byte-identical to the tree that had created claims.

The opening rule still wins: a change fully bound by a focused test on the real code path, and a question only a recorded live run can answer, get no steward — mark the second unobservable rather than manufacturing a finding.

A cleared deterministic condition is not finished until its owning test file carries it: the gate must bind the join the condition proved, or the next change reopens the question. The recurring shape is two tests that each hand-roll the same record — the watchdog readiness file on 2026-09-01 — so neither runs the real writer against the real reader. cases/layer-walk.md names the method.

A steward runs its condition on an account of its own. A steward subagent may own one condition again (operator decision 2026-09-29), but only when that condition's model calls go through a numbered account that no live run and not the launching session is spending. Without such an account, the launching session runs its own conditions, as it did from 2026-09-15.

  • Choose from .accounts/usage. Run it from the main checkout. It prints each numbered account's 5-hour and weekly windows and prints no token. It marks the account behind the plain CLAUDE_CODE_OAUTH_TOKEN, which is the one launch-run hands to live Claude runs by default. Pick an account that is:

    • not the plain one, and not the one a live run was launched on with .accounts/launch claudeN (scripts/credential-use.mts names the credential file each open run's launch receipt holds);
    • not the one the launching session runs on;
    • showing neither rejected nor a window near full.

    The reading on 2026-09-29: claude1 at 3 % (5-hour) and 6 % (week), claude2 as the plain token at 88 % of its week, claude4 rejected.

  • Pass the account as an env file. Run every model-calling script as bun --env-file=.accounts/claudeN.env .claude/skills/system-path-simulation/scripts/<script>.mts …. Every slot resolves its credential through loadRepoEnv(repoRoot, Bun.env), and there the process environment wins over the checkout's .env. So this one flag moves the Builder, Built, Judge and review slots together. A shell that already exports CLAUDE_CODE_OAUTH_TOKEN beats the file, so check that it is unset. run-condition.mts records it and names open runs sharing it.

  • The runners drop the launching session's variables. run-condition, run-segment, review-settle and judge-replay strip and print each CLAUDE* name the product does not read (CLAUDE_EFFORT, CLAUDECODE, …; session-env.mts), which it would hand on to the Builder's CLI. A paid launch starts under env -i and never has them.

  • Treat the files as secrets. .accounts/ is local and untracked (.git/info/exclude). Each claudeN.env holds that account's token under CLAUDE_CODE_OAUTH_TOKEN, the one name the launcher reads, and is written by .accounts/launch claudeN.

    • Never print, copy or cat these files, and never put a token on a command line.
    • If a file is missing, report it. Do not rebuild it from .env by hand or look for another.
  • The steward's own session still spends the parent's account. Only the condition moves, so keep the steward's reasoning short and let the condition's calls carry the spend.

  • The steward's shape. One steward per condition, at most four in parallel, and none spawns further agents. Each writes its trail to the report path it was given, so a dead steward loses nothing the parent cannot read. The steward briefs the parent with the account it used and that account's .accounts/usage line before and after.

The actor under test keeps the run's own condition — for the standard Opus run that is claude-opus-5-5 at medium on the Builder, Built and review slots, launch-run's opus preset, which run-condition.mts --preset opus pins — and neither the session nor a steward ever stands in for it (operator decision 2026-09-01). The unpinned Claude default in src/backends/resolve.ts is still claude-opus-5, so a runner that takes --backend claude without a preset serves a different actor.

Rules that hold for every case

If a gate or test already binds the thing you are about to check, you are finished before you start.

Spend alone is not the criterion (operator decision 2026-08-16, restated 2026-08-17: "we do not care about the money, only indirectly by checking if the model is doing sensible things; time is more important"). Use a live behavioural condition when it can change the decision and the host already has action authority. Keep the epistemic discipline: one stubbed thing, one variable, and stop at the decisive fact because work past that point confounds the answer.

Prefer a new discriminator to an idle continuation. A condition asks a new question; a further turn often repeats an answered one. Of four conditions run on 2026-08-17, the control changed the conclusion. This is evidence for good experimental judgement, not a required condition count.

Keep independent live-model conditions independent. Give each one exact head, one evidence packet, one question, one scratch directory and one report path, as its own background command, and keep every other model-visible byte identical. Check each load-bearing claim against the named bytes before routing it.

Seed as late as the question allows. The segment starts at the last stage whose output the question does not depend on; everything earlier is paid for and confounds nothing. Copy the required recorded source bytes into an owned fixture before invoking a mutating production helper. Inspect linked runtime and tool trees too: a copied workspace can still write through .toolchain into the original run, outside its fingerprint. Use owned tool fixtures for startup and resume probes; keep real recorded trees read-only. A republish opens a fresh epoch, so the source epoch's MEMORY.md and SCRATCHPAD.md do not follow it (production carries notes only when an epoch supersedes one); the seed names their sizes, and the stated delta must too.

Optional: give the condition an active tool tree. Do this when the actor will compile, run or rehearse through the product's tools. Without a tree it spends its wall reinstalling, and the condition ends up measuring installation instead of the question. Skip it when the question is prose or planning, or when a reinstall is cheap. On 2026-09-30 a truss Builder with no recorded tree reinstalled OpenSeesPy from the uv cache in 6 seconds; an esp32 tree is 12 GB of arduino cores.

  • Seeding already clones the tree. seed-campaign.mts republish clones the selected product's tool tree into the owned seed with APFS clonefile and rewrites the references inside it to the seed's own tree by rule, under every spelling of the source; --relocate now answers only references outside the tool tree.
  • Check the tree exists before relying on it. scripts/tool-tree.mts --campaign <dir> says whether the selected product's recorded toolTree still resolves and lists the family's other trees; every epoch tree of the truss campaign …-3fd52f9e-29 was swept on 2026-09-30. The seed and run-condition.mts both warn of a gone tree, --tool-tree <tree> seeds a family tree instead, or let the actor reinstall and record that it did.
  • Clone anything the actor may write; do not link it. A symlink into a recorded run lets the condition write through to that run, and the audit refuses it as an escape. Link only to a host file that no recorded run owns and the condition cannot change, such as ~/.bun/bin/bun.
  • Give each parallel condition its own seed. A campaign holds one .controller.lock, so two conditions or live arms from one position each need a seed; --as-slug a,b,c seeds them in one call. The esp32 ablation arms of 2026-09-30 were three seeds of one campaign.
  • Write it in the trail. seed.json records the tree's source (toolTreeSource); add its size and whether the actor reinstalled anyway, which is a finding about the seed, not the product.

Choose the production owner before the helper. A session measures authoring, a Builder campaign also measures continuation and admission, and a full run adds measurement and routing. Read the actual exports and helper flags at the measured revision; a familiar script name does not prove that it still mounts production tools or supplies the current read grant and system prompt. For admission or semantic comparisons, read authoring-comparison.

Fifteen turns is the maximum budget (operator decision 2026-08-19, replacing the ten-turn limit from 2026-08-17). This is the whole segment's model-turn ceiling, not a tool-call limit: one turn may contain many tool calls. A condition may request fewer turns but never more. A segment that cannot answer inside about fifteen turns is badly staged, not under-funded: seed later and ask something narrower. turn-budget-reached is a result to read, not a failure to retry. Keep any smaller authorised bound. Turns, tool calls, submissions and elapsed time are separate: one native turn can contain six submissions and hours of real compiler work. Record progress and settlement at their own boundaries; do not add turns, nudges or a new source revision to an active condition. An explicit wall limit is a separate part of the condition.

A settled actor ends the step. A continuation that calls no tool closes the step as step-settled. On 2026-08-17, 15 of 22 turns across four conditions were continuations that called nothing and answered "already complete", each still re-reading a transcript that had grown all segment — the idle turn sent 49% more input than the working turn before it. A stage with more to do needs a continuation carrying a reason (--continue-file), not an empty nudge. Spend reads as a signal here rather than a limit: a condition whose turns cost a lot while its trail shows no tool calls says the actor stopped working.

Read the results before starting anything else. The recorded failure of this script is an unread answer rather than a wrong one: four conditions finished while their pre-registered predictions sat unresolved. The script closes with an UNRESOLVED line naming each P<n>. Resolve them against the trail and the workspace, and append each resolution with scripts/predictions.mts --resolve; --unresolved lists what is still open.

Scout every prediction to its bytes before the actor starts. For each P<n>, name three things: the surface the chosen runner mounts that can produce the observable, the recorded file where it lands, and a reader already shown to see it on one known positive. If any is missing, the condition cannot prove the prediction however long it runs. On 2026-09-15 a census condition went through run-segment.mts, which mounts no gate, and a status probe counted a process name the suite never started. It ran 78 minutes and 70 authoring calls, then resolved all three predictions untriggered. run-segment.mts now refuses predictions naming a gate, census, F2, adoption or battery before any backend opens (--accept-unreachable "<reason>" keeps a deliberate absence question); route those questions to run-condition.mts. A monitor's first tick should find a positive, not only zeros.

Pick your case, then read it

Classify the question against the table, read the case file (or files) under cases/, and follow it. One question usually needs one case; a batch of conditions may need several, read in the order the table lists them, because the cheaper case often cancels the dearer one. Every case uses the position rules below.

The question looks likeCase file
One fact: which branch, value, identity, which version the wrapper runscases/check-one-fact.md
Will an actor follow this instruction — and can the stack under it carry it?cases/layer-walk.md (always before a live condition)
Does this path hold end to end with one stubbed step?cases/scenario-stub.md
What does the real model do from this position, in a few turns?cases/live-segment.md
Can the Builder reach admission, repair a semantic gap, or improve under another model condition?cases/authoring-comparison.md
Seed a recorded campaign and run one real controller round with one slot live, the rest scripted or offcases/seeded-condition.md
Does the whole runtime hold: controller, wall, verifier, record?cases/seeded-project.md
Compare two trees or conditions on whole runs; rehearse a prompt on a full runcases/fullrun-conditions.md
Does a fresh project on this head reach its first preview, adoption and the Built Harness?cases/real-run-watch.md
Two or three past runs are named: which to analyse, and how to re-run it herecases/past-run-replay.md
What the live Judge or Epoch Reviewer now decides from a recorded casecases/review-replay.md
"The brittle angles", "what would break", across several runscases/brittle-angles.md
A final test of a large PR stack: which single PR carries each errorcases/stacked-prefix-groups.md

A past run beats an authored position. Before writing a situation by hand, run scripts/pick-run.mts over notes/runs/ and take the nearest recorded one; the position section below says how to derive the rest from its recorded bytes.

Derive the position, and keep your hands off the outcome

Every shape needs a position: where the actor stands, which bytes it holds, what already happened. The same hand writes the position, the conditions and the verdict, so this is where a simulation stops measuring the product and starts measuring its own prose.

How much of the position is real

A position is four separable facts, each derived or authored on its own:

  • the tree — campaigns/<name>/epoch-*/workspace/, straight into --seed-dir;
  • the model-visible text — the contract and framingDigest in epoch-*/builder-session.json, so a seeded kickoff is proved byte-equal rather than asserted;
  • the controller state that put the actor here — controller/<fullrun-*>/opening.json and terminal.json, plus the recorded case rows for that checkpoint;
  • the history the actor is told it lived — the misleading part. A Claude slot's own transcript lives in $TMPDIR/ana-claude-cli-*/projects/<cwd-slug>/*.jsonl only while its session is open; scripts/condition-evidence.mts keeps each state of it during a live condition, and a past run's epoch-*/builder-prose*.jsonl keeps only the words. scripts/position-packet.mts copies a transcript's last three to five exchanges verbatim — tool calls and results — under a summary labelled as authored, so only the summary is yours. Ten exchanges drown the situation (operator decision 2026-08-23). A Codex slot leaves a rollout under ~/.codex/sessions in another shape; there the history stays authored and labelled.

Deriving three and labelling the fourth beats a whole position hand-built for coherence. Label them where the predictions are written: a finding is as strong as the weakest part the behaviour depended on.

Usually no run stood exactly here, and then the quantity is the delta to the nearest real position: one battery earlier, eight tasks instead of 25, a rebuild advice packet the real run had settled. A delta you cannot state in one line means the position was decorated rather than derived. The condition is part of it — four conditions on 2026-08-17 ran at --effort medium and low against a run that pins higher, which is a different actor rather than a cheaper one.

Two traps run the other way. Real bytes are not automatically the right bytes: seeding from a run that predates the change under test measures the old product. And a position no completed run has reached raises reachability before behaviour — one recorded pair spent 910 and 431 seconds landing twice on the same verifier-required terminal, answering reachability by accident.

When the block and the bytes disagree, the behaviour is a response to the story. The climb condition of 2026-08-17 declared a frozen adopted climb round over a seed its own predictions file records as the w28 workspace "standing in as" one, shared with three other conditions.

Whether the position hands over the answer

Production puts an actor in a situation. A seeded position tends to hand it an assignment instead, and the recorded uses show four ways that happened:

  • An imperative production never sends. "Decide and take your next action now" closed all four Builder positions of 2026-08-17; "then stop: report what you set up" closed the toolchain condition of 2026-08-16.
  • A named target, which turns a propensity question into a capability one. One condition asked whether the Builder can reach the network from its walled workspace, and its position told it to install PlatformIO. "Can, when told" and "does, unprompted" are different findings, and a run usually depends on the second.
  • The falsifier stated in the subject's own prompt. "One more byte-identical submit ends this session" stood in the position of the condition predicting the session would stop resubmitting.
  • A prediction whose caught is compliance. P1 and P7 that day both resolve caught when the session does what the steering text quoted in the same prompt asked; neither condition could fail. P2 is the shape that pays — it asked whether the edit landed on the surface the finding named, which the prompt did not say.

The line that holds: production constrains scope, it does not assign the next action. A position may name the files writable this round, because production does. It may not say what to do with them, in which order, or when to stop.

Steer by transition, not destination

The useful behavioural shape is A → B → {C, D}:

  • A is the actor's settled past, derived from recorded bytes: what it built, what was measured and what decision boundary it reached.
  • B is the new controller-owned situation: the selector moved, a scope opened, or a public fact became available. Append B at the boundary where production would deliver it, normally through the real handover rather than by putting future history in the first kickoff.
  • C and D are plausible outcomes. They belong in the campaign-owned frozen prediction/event ledger and its advisory simulation projection, never in model-visible text. The expected or "right" answer must not appear as a named target, suggested mechanism, falsifier or stopping instruction.

Write the transition as ordinary continuity: "You were in A. The controller has now selected B. The workspace is exactly as that completed stage left it, and this round has scope S." This gives the actor enough orientation to proceed without assigning the action it should take. Use the real controller's reason and contract where they exist; an authored transition is labelled as such.

For an authored A or B, pre-register and run three conditions when the question is important enough to change production prose:

  1. Natural transition: A followed by B, with neither C nor D named.
  2. Context-minus control: the same bytes and contract without the authored transition sentence.
  3. Counterfactual transition: the smallest legitimate B′ that should change the correct branch.

The first pair measures whether the transition steered behaviour. The counterfactual measures whether the actor read its meaning rather than merely following its vocabulary. If B and B′ lead to the same branch, do not claim contextual reasoning. Keep all other model-visible text, model pins and tool contracts identical, and record each prompt digest separately.

Re-reading your own position for this does not work, since you wrote the sentence. Two checks instead: every imperative in the position must be quotable from a production surface, and one condition runs with the sentence removed — a behaviour that survives removal belongs to the product.

Say what you did not run

Pre-register the prediction rows before the first condition. Condition selection, dropped conditions and reruns belong to the campaign-owned frozen prediction/event ledger; this simulation note is only an advisory projection of those rows. predictions.mts --hash writes a local integrity checksum for the projection and refuses to replace it after the note changes. It does not create campaign authority, consume an allowance, choose an experiment or promote a result. A rerun or newly justified condition gets a new campaign prediction/event, not a silent edit to this projection. State the expected outcome and what would refute it. Include an expected-failure condition when the evidence supports one; a demonstration with no discriminator is not an experiment.

Scripts

All under .claude/skills/system-path-simulation/scripts/; each takes absolute paths and refuses a relative path or an unknown option.

ScriptDuty
pick-run.mtslist recorded runs from notes/runs/ with ancestry, denominator and component facts for comparison
position-packet.mtsthe last 1–5 exchanges of the actor's own SDK transcript, verbatim under a labelled authored summary, with the source digest
difficulty-watch.mtsthe authoring sequence and submit rows from recorded builder-execution records, codes only
predictions.mts--hash, --verify, --resolve, --unresolved on the prediction note
run-segment.mts, seed-kickoff.mtsa seeded live segment over the production backend
workspace-changes.mtsa workspace's changed paths and diffSha, counted from its seeding commit, since checkpoint commits move HEAD (--since HEAD for the view since the last one)
seed-campaign.mtsclone a recorded campaign into a fresh tree, or republish its selected product here under one or more new slugs (--as-slug a,b); symlink and absolute-path audit, the owned tool tree relocated by rule, seed.json; names what it could not carry (a swept tree, the epoch's notes), and --tool-tree stages another tree
tool-tree.mtswhether the selected product's recorded tool tree still resolves, its size, and the family's existing trees; --digest says which match the recorded digest
run-condition.mtsone real controller round over a seeded slug with each slot live, a scripted module, no-solve or capture; --preset as launch-run pins it; --capture compares a first prompt with an earlier capture; preflight of session variables, Built pin, tool tree and credential; wall, sampled process census, preregistration digest, report.json
condition-evidence.mtsa condition's live CLI transcripts and durable Builder records, kept under content names that never overwrite (--follow while live), a per-session tool tally and --grep over the prose
credential-use.mtsthe credential file each open launched run holds, from its launch receipt; names only
judge-replay.mtsthe live Main Judge over recorded battery cases under the current prompt; verdicts and every contested row
review-settle.mtsthe live Epoch Reviewer over a scratch copy of a recorded battery and its Judge disagreements; the evidence and its public projection
host-panel.mtsvalid, equivalent and hostile artifacts for one task through the real verifier host; fingerprint before and after
show-prompt-surfaces.mtsthe exact model-visible surfaces and their digests

The build stage has one interface, HarnessBuildOptions.builderRuntime (src/run/harness-build.ts). Production binds the live transport; a scripted session bound there still runs the real submit, census, solvability and adoption path. test/full-run-scripted-loop.test.ts runs the whole controller loop through it with no provider, and test/helpers/scripted-builder-runtime.ts is the one scripted session both that warranty and run-condition.mts use. Do not hand-build a session, a seed clone, a product pointer, a wall, a process census or a host panel again: each has an owner above, and the recorded setup faults of 8 to 10 September all lived in hand-written copies of them.

Separate your script's faults from the product's

A simulation produces different kinds of refusal that can look identical on stdout. Classify the owner before attributing behaviour; repair setup faults and verify the repaired path within the existing authority before making a product claim. Recorded artifacts, all from simulations that were otherwise sound: a scratch tree with no node_modules above it failing as Cannot find module '@ana/agent-bundle'; a helper typed level: number | null = TARGET_LEVEL, so the "level omitted" scenario passed undefined and triggered the default it claimed to omit; a script that never passed expectedTasks, making its own silence read as a battery shortfall; a scratch copy that symlinked node_modules, which fullrun refused before any provider work; a probe that called the wrong export and failed before reaching the product fact at all; a seed exported with git archive and therefore without .git, which initWorkspace overlaid with the starter skeleton so both conditions rebuilt from placeholders while the note declared an adopted-product position (2026-09-13; run-segment.mts now commits such a seed as the root and verifies the bytes).

For every refusal you intend to report, name the production caller and confirm it passes the same arguments your script passed, from a tree set up the way production sets one up. If the real caller passes more, your scenario is not yet the production scenario. Preserve the failed attempt and classify its owner: simulation setup, shared product code, generated candidate, or provider/host environment. A repaired setup gets a new labelled attempt; do not fold it into a model's failure count. Reproduce a shared defect against unchanged candidate bytes when possible, then add the positive and nearest hostile case to its existing owning test.

Test the helpers without model spend

Run a changed script's owning suites through the prepared-worktree runner; none calls a model:

bun run test -- test/sps-*.test.ts test/system-path-simulation-*.test.ts test/full-run-scripted-loop.test.ts

One test/sps-<helper>.test.ts owns each helper, and test/sps-refusals.test.ts holds one row per declared option refusal: select only the ones owning the change. Bun never descends into this hidden directory, so test/test-discovery-completeness.test.ts refuses a suite written beside its helper. Zero discovered tests is not a pass. A prose-only skill edit needs frontmatter/link validation and diff review, not this runtime matrix or a paid model replay. A passing matrix is mechanism evidence for the helpers, not for the model behaviour a later condition measures.

Result

Return one evidence path and one result:

  • blocking: the command would use the wrong branch or condition;
  • cleared: the check proves the intended path;
  • accepted-risk: only a paid run can decide it.

Return this result to the invoking agent. cleared preserves any existing authority but does not exercise it, schedule a command or choose the next experiment. The improvement loop or standalone host owns that decision and launch boundary.

A deterministic check never predicts provider behaviour, output quality or later runtime branches; for those, accept the risk or run the rehearsal in cases/fullrun-conditions.md, whose result is the prediction note with its resolution section, not one of the three words. For live conditions also report operational closure, semantic verdicts and prediction resolution separately. Preserve verified, unaccepted and non-result denominators; list unexecuted probes. A host-blocked condition can leave a prediction unresolved. Admission, a compiler exit and a passing generated control set each prove their own condition, not complete domain correctness.

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

同じリポジトリのスキル

概要と使いどころ

Use when deciding which Anabasis stage owns a change, when carrying one bounded build/measure slice from the user's prompt to evidence, or when deciding what may enter a run: the exact one-line prompt, context files, public catalogues, research and solve-side public data. Maps the product loop, the input contract, the authoring sessions, the adoption gates, the four owners, and validate-or-reopen behaviour. Not a long-running campaign controller.

日本語の概要は準備中です。原文の説明を表示しています。

s-smits/anabasis612026年10月8日 更新

Use after an Anabasis run, comparison, or system change and before claiming improvement. Attributes movement to one owner, reports identities and censored denominators, separates deterministic proof from model judgement, keeps measurement validity apart from observed exploitation, and states what remains unrun or provisional.

日本語の概要は準備中です。原文の説明を表示しています。

s-smits/anabasis612026年10月8日 更新

Use for 1-7 independent read-only sessions investigating one bounded Anabasis failure, diff, or design question before a fix. Defines non-leading prompts, evidence packets, symmetric fault hypotheses, self-falsification, source adjudication, and concise synthesis. Do not use for whole-run coverage at any session count; use whole-run-investigation instead.

日本語の概要は準備中です。原文の説明を表示しています。

s-smits/anabasis612026年10月8日 更新

Get told the moment a live Builder session is blocked (five refused submits in a row, a repeated findings set, twelve checks without acceptance, two environment previews in a row, two hours since a submit or rehearsal last returned) and fix the owner iteratively: read, fix on the owning PR, then keep the run or kill it and relaunch on the upgraded system, and rearm. Use while a paid run is live and the operator wants blocking caught and repaired, not only reported.

日本語の概要は準備中です。原文の説明を表示しています。

s-smits/anabasis612026年10月8日 更新

Shared vocabulary for designing deep modules. Use when the user wants to design or improve a module's interface, find deepening opportunities, decide where a seam goes, make code more testable or AI-navigable, or when another skill needs the deep-module vocabulary.

日本語の概要は準備中です。原文の説明を表示しています。

s-smits/anabasis612026年10月8日 更新

Launch, start, monitor, and collect independent Codex subagents (gpt-6-luna at high, xhigh or max; gpt-6.1-sol for small review batches) for bounded parallel work, from Codex or from Claude Code, including requests supplied as a Markdown file of session prompts. This skill owns Codex subagent transport even when another investigation or review skill defines the questions. Use whenever the user asks for Luna or Sol agents, a swarm, many sessions, a concurrency test, a particular reasoning effort, a Codex subagent from Claude Code, or later collection of reports. Distinguish launch-only requests from requests to wait, collect, or synthesise. Route high and xhigh through the direct launcher; for 16 or more sessions, also invoke the direct launcher instead of native spawn_agent.

日本語の概要は準備中です。原文の説明を表示しています。

s-smits/anabasis612026年10月8日 更新

s-smits のスキルをすべて見る

このスキルの問題を報告する