aax
無料Optional AAX support for Pulp, including developer-supplied Avid SDK setup, CMake enablement, DigiShell/AAX Validator workflows, and local AAX builds on macOS or Windows.
日本語の概要は準備中です。原文の説明を表示しています。
Forge Modular's generator, patch checker, module pack and Forge-worktree seam — the traps that make green results untrue
インストール方法を見るインストールする前に、エージェントに与えられる指示の中身を確認できます。
Covers tools/rack/ (the generator, the patch idiom library, the gates),
examples/forge-modular/ (the module pack), and forge-seam/ (the app's
sources, which live here and are built inside a Forge worktree).
Read this before trusting any green result in this area. Everything below cost a session at least once, and most of it was a test or a tool reporting success for work it had not done.
core/simd/include listedSignal headers (fir_filter.hpp, oversampling*.hpp, zero_latency_convolver.hpp,
and everything that includes them) include <pulp/simd/simd.hpp>. A build
that lists include directories by hand must list core/simd/include next to
core/signal/include, or it fails with a missing header. It needs no library:
without pulp-simd's compile definitions the header supplies the scalar kernels
inline (pulp::simd::active_backend_name == "inline-scalar"). Four lists carry it
for this pack and must move together: examples/forge-modular/CMakeLists.txt
(both include blocks), the consumed-input status list in package.sh, and
the toolchain tree list in tools/rack/install_toolchain.sh, and the generator's
INCLUDES in tools/rack/generate.py (preflight and syntax compilation). A pack that
misses it fails at the compiler after the model has already been called.
Linux refuses any single argument over 128 KiB and macOS refuses argv plus the
environment over 1 MiB, and both fail with E2BIG before the tool starts, so a
long request cannot travel on the generator's command line. patch.py build
and generate.py therefore accept --prompt-file PATH in the prompt's place:
prompt_handoff.resolve() puts the file's text back into argv at that position
(decoded as CPython decodes argv), deletes the file, and refuses a NUL byte. A
positional prompt still works from the command line.
The app half lives in the Forge repository, not here: Forge's
src/process_engine.cpp is the shipping ProcessEngine, and Pulp's
forge-seam/modular/ copy is an older staging snapshot that no build uses.
Change the engine in Forge, and test the app-to-generator boundary there
against these tools. A stub generator in a Forge test must read
--prompt-file, not sys.argv[2], once the engine passes it.
patch.py build and generate.py take generation_lock.acquire_or_exit() on
the Rack plugin directory they install into (RACK_PLUGIN_DIR, or
generate.py --install-dir) before any work, and exit 75 with a plain reason
when another generation holds it. The lock is an flock on
<lock root>/<sha256(realpath(dir))[:32]>.lock, never a marker file and
never inside the plugin directory. The kernel releases it when the process
dies, crash or SIGKILL included. The lock root is FORGE_GENERATION_LOCK_DIR,
else ~/Library/Caches/Forge Modular/generation-locks on macOS. The Forge
engine must compute the same path (realpath, sha256, first 32 hex) to take the
same lock. Tests set both variables to temp directories, so they never touch a
real plugin directory or contend with a running Forge.
generate() returns (patch, why, shortfall). shortfall is None on a pass
and a Shortfall when the loop ran out of attempts and is handing over the
best thing it built. It only raises SystemExit when there is genuinely
nothing to hand over — no attempt ever got past the lint.
The rule this encodes: an advisory verdict is not fatal, and out of attempts
is not the same as nothing to show. The loop used to raise SystemExit("gave up after 5 attempts") after generating five patches its own transcript had
called "structurally sound" and "keeping the patch anyway", and the user got
nothing — not the patches, not a copyable reason. "Not a sequenced-voice patch
yet" is a wording that means the loop should keep going, not that the patch
is worthless.
Consequences to keep in step when touching this:
<slug>-unfinished.vcv, prints the "built N
modules, M cables → path" line the app parses for an artifact, then the
handover_report() block, and exits 1. Dropping either the exit code or
the gave up after phrase breaks a different reader: the shell decides a run
ended by classifying that phrase, so without it the app watches a dead build
forever.Shortfall. It names modules
Rack cannot create, so handing it over offers a file that will not open.Shortfall.rank is (severity, misses, -attempt), lowest wins: a patch that
plays beats one measured silent, fewer missed requirements beats more, and
only a tie goes to the later attempt. Rank on requirements missed, never on
the number of LINES it takes to say so — naming installed jacks adds lines,
and ranking on those would rate a patch worse for every jack we managed to
name.keep_attempt is not gated on FORGE_ATTEMPT_DIR any more. It was, and
nothing but prove_idioms.sh ever set it, so every app build kept nothing
while the log claimed otherwise. The default is a stamped per-pid directory
under ~/Library/Application Support/Forge Modular/attempts/.name_the_jacks() turns each failed requirement into something the next
attempt can act on. The requirement itself is a sentence the model already
believed when it wrote the patch — "the sequencer's gate has to fire an
envelope, or every step runs together" — so handing it back as the correction
produced five near-identical attempts and nothing escalated.
It joins the missing describe-strings back to the idiom's topology entries
(the strings ARE req["describe"]), then adds up to three lines per side:
this patch's sequencer CANNOT send it, however it is wired: CVfunk/PentaSequencer (its outputs are A, B, C, D, E) — the most actionable
line there is, and the one that ends the loop of trying the same module.installed jacks that can send it: CVfunk/StepWave out1 'Sequencer Gate'Rules it follows, each of which was a wrong answer first:
any_out/any_in) names nothing, because
every jack matches and the list would be noise.Two lessons from the run kept in tools/rack/test_fixtures/silent-oscillator/.
A correctly-wired patch read silent because one oscillator reported 0.000 V,
and the model kept that oscillator for four consecutive attempts, adjusting
its knobs. Its params were all in range and none at a silencing zero, so
retuning was never going to work.
That oscillator was not broken. It was licensed, and the gate of the day never
resolved licence keys, so it wrote zeros: this run is an instance of "A
licensed module runs, writes zeros, and logs nothing" below, and the replay in
tools/rack/licence-fix-replay.json carries both of its patches moving silent
to audible under the fixed gate. That changes the cause and not one of the
lessons. The retry could not see the harness, only the readings, so naming
the first module in the chain reading zero was still the right move and the
right answer was still Zephyr.
The information was already there and was not usable. The retry context
carried the entire per-module table and told the model to "find the FIRST
module in the chain whose output is 0.000". It did not. Handing over data plus
an instruction to infer the conclusion is not the same as handing over the
conclusion. dead_module() names it: CVfunkSands/Zephyr PRODUCED NOTHING: out0=0.000, out1=0.000.
Rules that finder follows, each of which was a wrong answer available:
PressedDuck also
reads 0.000 and is what the FAIL line points at — it is silent because
Zephyr is. The cause is a dead module with no dead module feeding it, decided
from the patch's own cables, never from the order the gate printed.(not instantiated) reported nothing, which is not zero. Core's audio
interface is in that state in every patch, and calling it dead would name the
one module that cannot be the cause.The more valuable half is escalation. Each attempt arrives knowing only its
own rejection, so repetition is a fact only the loop can see — and it was
telling nobody. The fourth prompt read exactly like the first. generate()
keeps silent_runs and missed_runs; from the second consecutive attempt with
the same dead module (silence_advice) or the same unmet requirement
(stuck_note) it says to replace the module, not retune or rewire it, and
names installed candidates. On the real evidence this fires on attempt 2, three
attempts before the model found its own way out.
A different dead module, or a different requirement, is progress and must never be called a repeat: escalating on a run that is improving tells the model to throw away a fix that worked.
Do not send another model call when the full Rack gate proves that a patch's
listener path is silent but the source module has a measured-live sibling
output. deterministic_repair.py may audition that sibling only when its exact
semantic role set matches the silent output, and it may change only cables on
the listener path. Activity is not acceptance: rerun the full DSP/audibility
gate, then apply every existing blocking prompt-derived static, runtime, and
long-horizon qualification to that exact candidate. Continue through the
same-role siblings until one is both audible and fully qualified; audibility
alone must not short-circuit a later sibling whose real DSP satisfies the
request. If none fully qualifies, retain the audible candidate with the fewest
missed requirements as unfinished with its exact errors; preserve the original
patch only when no sibling became measured audible. This is a general
output-selection repair, never a module-identity exception.
The acceptance result has to remain bound to the candidate that was actually
measured. patch.py may not reuse an earlier long-horizon result after a
deterministic repair changes the patch, and a timeout or unavailable listener
is UNMEASURED, never a pass. When the gate prints a completed artifact or a
specific terminal failure, drive_app.py must classify that exact producer
wording as PASS or FAIL rather than degrading it to INCONCLUSIVE. Keep producer
and every reader pinned together with positive and negative fixtures.
render_inventory has the same honesty rule: a module with no recorded jacks
prints ports: UNKNOWN, and inputs and outputs are independent. It used to
emit both only if m.get("inputs"), so an uncartographed module and one with
no ports at all rendered identically, and a module with outputs and no recorded
inputs lost its outputs.
Both gates are standalone binaries that take a patch and a plugin directory, so replaying a failure is free where reproducing it costs a dozen model calls:
PATCH_GATE_TRACE=1 DYLD_LIBRARY_PATH="$HOME/SDKs/Rack-SDK" \
~/.cache/forge-modular/patch-gate <patch.vcv> \
"$HOME/Library/Application Support/Rack2/plugins-mac-arm64"
prove_idioms.sh keeps every failed attempt's patch and full report under
tools/rack/patch_idioms/regressions/idiom-proof-logs/NN-attempts/. Three
separate defects were found and fixed there in the time one model call takes.
If a failure is not diagnosable from what is on disk, fix the harness first —
that is cheaper than another run, every time.
.vcv is not the JSON Forge originally wroteForge emits a new patch as plain JSON, but Rack 2 saves that patch back as a
Zstandard-compressed tar archive containing ./ and ./patch.json. Reading a
Rack-saved file with json.load() therefore reports a misleading UTF-8 error;
moving or renaming the file is unrelated.
rack_open.py distinguishes the two formats from the immutable input bytes.
Archived patches are decoded only by the adjacent rack_patch_decode helper,
built from the pinned vendored zstd decoder. The helper parses the tar in
memory, accepts exactly one regular root patch.json plus Rack's zero-byte root
directory entry, caps both input and expanded data, and rejects traversal,
links, nested directories, duplicates, bad checksums and trailing ambiguity.
It never extracts a member and never depends at runtime on Homebrew, tar,
zstd, Python extensions, or Rack itself.
Build the helper on the maintainer/package path with
tools/rack/build_rack_patch_decode.sh build/rack_patch_decode. The signed app
ships that prebuilt binary beside shape_text, and install_toolchain.sh only
copies and validates it; an ordinary end-user install must never require a C or
C++ compiler. Cross-architecture packaging sets
PULP_RACK_PATCH_DECODE_ARCH and verifies the helper's exact Mach-O slice;
running a helper on the packaging host does not prove it can run on the target,
and a foreign target slice must not require Rosetta merely to verify a package.
Do not load libRack.dylib through Python ctypes to decode this format.
Rack's library expects application runtime initialization; an exploratory call
reached OpenSSL BIO_new_ex with invalid state and crashed Python with
EXC_BAD_ACCESS. Keep the crash boundary out of the app process and preserve
the packaged-helper tests under a minimal Finder-style PATH.
patch_lang.render() names each module from its model slug and numbers a
repeat. Appending the count alone collided: a second VCO became vco2, the
name a VCO2 already has, and the parser refused its own output ("'vco2' is
declared twice"), so a library patch with two VCOs and a VCO2 silently stopped
round-tripping. Every name is now checked against the names already given, and
a slug ending in a digit is numbered after an underscore (vco2_2). Keep that
check if the naming changes again.
test_patch_lang.py reads tools/rack/fixtures/patch-lang/inventory.json, a
checked-in slice of the inventory covering every module it, the module
docstring and docs/contracts/patch-language-v1.md declare, so it runs in CI
with no Rack or Forge data. A new example that declares another module must
add that module's ports and panel there, or the test fails on the example. The
round trip over the installed example patches still runs against the live
inventory where they exist and is a printed SKIP elsewhere.
out0=... is the instantaneous voltage at the end of the run. The
silence verdict is a separate figure: mean_abs over the whole run, and only
for cables into the audio interface. A module showing 0.000 in the trace is
not necessarily silent, and vice versa.ch= is outputs[0].channels — the FIRST output's channel count, not the
module's. It is a red herring nine times out of ten; a whole diagnosis was
built on it once and had to be thrown away.| pN=value after the pipe is the params. Check them before the wiring.
A level, gain or amount knob MULTIPLIES its CV, so p0=0.000 on a VCA is
silence however hard its envelope fires — and the gate's advice used to say
"its CV never rose", which sent five attempts to re-trigger an envelope that
had been firing the whole time.sounds() no longer existspatch.py::audibility(patch) returns (AUDIBLE | SILENT | UNMEASURED, report).
It was sounds() -> (bool, str), and the rename is deliberate: a caller written
against the boolean now fails loudly rather than reading a non-empty string as
success.
UNMEASURED is the whole point. A missing SDK, an unbuilt gate, a gate that
died, a run that never finished, and a gate refusing its own PATCH_GATE_SET
are all "this patch was never judged" — not "it passed", and not "it is
silent". The boolean forced that choice and got it wrong: with no SDK it
returned True and a run printed audibility as passed having measured nothing.
That is the same defect as a gate that measures presence, one layer up — the
absence of a failure reading as the presence of a pass.
generate() keeps an UNMEASURED patch and says the doubt out loud, because a
patch that lints clean and whose audibility is unknown is worth more than no
patch. It used to decide that by sniffing GATE_CRASHED out of the report's
wording, which covered the crash and nothing else — so the no-SDK case took the
other branch and printed nothing. The verdict carries it now; don't reintroduce
a string match.
Audibility and behaviour are independent layers. held.vcv — a bare VCO
into the interface — is genuinely AUDIBLE and genuinely fails melodic. Both
are correct. Presence was never the property; it is also not nothing.
patch-gate runs 6 s of the real DSP and prints TWO things about every cable
into the audio interface: the old presence verdict (mean_abs, pass/warn/fail)
and a behaviour block that is measurement only — pitch variety, attacks,
level over time, brightness over time. Nothing in the behaviour block can fail
the gate. Whether a patch's numbers are the right numbers is a question about
what was ASKED for, and the request is not in that process.
patch_behaviour.py, which maps an idiom's behaviour
flags to predicates over those numbers. Every threshold is data in
patch_behaviour_thresholds.json and every one is an untuned guess —
first cuts, not measured truths. When a threshold disagrees with somebody who
listened, the listener is right; edit the JSON.UNMEASURED is not fail. A flag whose needs are unmet (a patch whose
pitch was trackable for 8% of the run) comes back unmeasured, because
rejecting on a number that was never really measured teaches the model to
satisfy the number instead of the request.voiced_fraction before believing distinct_pitches. A staccato
patch whose notes are shorter than min_note_windows (2 windows = 100 ms)
reports far fewer pitches than it plays. That is the measurement being
conservative, not the patch being wrong; PATCH_GATE_SET=min_note_windows=1
re-runs it looser, and the report records which setting produced it.periodicity is not reported at all below two onsets. A held tone's
detection function is flat to within arithmetic noise, and flat noise
autocorrelates at 0.99 — so a drone used to announce a confident pulse at a
meaningless lag.patch_behaviour.hpp is standard-library only, and that is load-bearing.
The gate stays a one-file clang++ compile against the Rack SDK and nothing
else, which keeps the SDK the ONLY thing standing between it and a machine
that can build it. It also detaches the analysis from the capture: the f0
estimator (cumulative mean normalized difference) and the radix-2 FFT are
written out here rather than pulled from pulp/signal, so
test_patch_behaviour.cpp builds with clang++ -I tools/rack and no SDK,
no Rack and no Pulp target. That half is gateable anywhere today. Do not
add an include path to this header — a -I or a -framework in either
build line means the property is gone, and check_behaviour_is_measured
fails when it happens.
(Linking pulp/signal in was tried and reverted; both estimators produced
numbers identical to Pulp's YinTrackerT and FftT on all eight synthetic
signals and both real patches, so nothing was traded for it. The Rack SDK is
GPLv3 — see tools/cmake/PulpRack.cmake — and while the shipped .vcvplugin
deliberately does combine MIT pulp/signal with it under Rack's plugin
exception, this binary has no reason to need the boundary considered at all.).cpp edit. GATE_HEADERS
feeds the staleness check; without it, editing patch_behaviour.hpp silently
reran the previous binary and the change appeared to do nothing.build_gate() returns (binary, reason). A compile failure used to reach
the user as "no SDK", which sent the diagnosis to the wrong place.PATCH_GATE_SERIES=1 adds the per-window tracks (semitones, RMS, centroid,
onset times) to the JSON. They are the whole diagnosis when a number reads
wrong and dead weight when it does not, so they are off by default.test_patch_behaviour.cpp is the proof, and it needs no Rack: it synthesises a
held tone, a five-note line, steady/jittered/accelerating pulse trains, a filter
sweep, a decaying tone and silence, and asserts each measures as itself.
--prove additionally runs every signal's expectations against every other and
fails when an expectation set describes nothing in particular — which is how
a suite that measures nothing goes green. It caught one on its first run.
An explicit #Tag and a claimed idiom are two independent constraints. The
idiom compiler narrows the model inventory to port-complete structural choices;
before that subset is rendered, intent_context.add_required_candidates() must
add fresh-generation-safe choices for every explicit tag. Otherwise the prompt
can say both "required" and "no installed module carries this tag", while final
lint still rejects the missing tag. That contradiction spends a provider call
on a patch that cannot pass. An impossible combination fails before the call;
never weaken the final tag check to make a narrowed prompt agree with itself.
An exact qualified module identity in a request is an output contract, not
only an inventory hint. After parsing a model response, patch.py must require
every positively named Plugin/Model identity to appear in the patch before it
can be saved or installed. Exclusions, replacements, coordinated removals, and
mixed clauses are classified separately; refinements use the same classification
for presence and instance-count checks. Prove this offline with
--response-file, including a retained response that omits the requested
identity and an identity-only mutation control, before spending another model
call.
Some Rack modules depend on opaque state stored in a saved patch that cannot be
reconstructed from the public parameter and port inventory. Record these as an
exact plugin-version fresh_generation = unsupported contract in
tools/rack/module-state-overrides.json. Fresh generation must hide them from
the model and refuse an explicit request before resolving or calling a provider;
refining an existing validated patch may still describe and retain them. Never
carry an exclusion to a different plugin version without fresh evidence. This
is why PathSet-Infinity/WarpDrive 2.1.0 is excluded: a newly authored instance
is silent without its opaque coil, LFO, and envelope sequences even though the
module is installed and its ordinary Rack metadata looks sufficient.
Geodesics-Vultiverse/Hexaquark 2.0.4 is excluded for the same class by
different evidence. It requires an external clock, and its musical sequence is
stored in an opaque roughly 32 KiB data.state; setting the visible Run or
Scene parameters does not author that state. Keep it available when refining a
validated saved patch, but do not offer or accept it for fresh generation until
an exact-version state compiler has a three-arm proof: known preset state plus
clock plays, the same state without clock does not, and blank state plus clock
does not. A newer version needs fresh evidence rather than inheriting this row.
Before a module-generation request can reach a model, it must select at least
one direct primary DSP capability from
tools/rack/knowledge/module/dsp-primitives.json. The installed public
capability manifest is necessary but not sufficient: a type can be exported by
the SDK and still be unavailable to Forge when this curated projection omits
it. The intentional symptom is this request matches only generic module helpers and no direct DSP capability; no provider request has started.
When adding an SDK primitive that Forge should author directly, add a narrow
primary row with unambiguous intent terms and a focused
tools/test_dsp_vocabulary.py positive. Also retain helper-only and generic
lookalike negatives so ordinary words do not accidentally authorize a model
request. Diagnose this projection first; do not retry the GUI or spend another
model call on the same prompt.
render_inventory() in patch.py is the catalogue the model receives. It
carries ports AND params (name, range, default). It did not carry params for a
long time, and the model wrote values blindly — which is how a kick drum came
back with its VCA level at zero. If you add a field to a manifest that the
model should reason about, it has to reach render_inventory or it may as well
not exist.
In a long-lived DAW host, never cache Rack's plugin directories at Python
import time. The directory may be created by the bundled-pack installer or a
Rack launch after the generator was loaded. Resolve it at the point of use;
when RACK_PLUGIN_DIR is set, treat that one sandbox-visible directory as the
authoritative install and scan destination.
“Only” or “exclusively” naming a maker is an output contract, not merely prompt copy. Preflight must prove at least one named-maker module is installed, and the retained patch may contain only those maker plugin slugs plus Rack's Core infrastructure. Reporting a substitution does not satisfy an exclusive request.
Two of these cost three attempts each, and both produce an EMPTY frame that passes every assertion around it:
MentionOverlay::build() returns a view that
positions itself absolutely, which resolves against a containing view. Render
it as a root and nothing is placed. Wrap it in a View, size the wrapper,
layout_children(), render the wrapper.build() comes before the content. Rows are made into the view build()
returns, so filling the list first leaves them nowhere to go.So any render test here needs a floor — REQUIRE(read_all(path).size() > N) —
or "it rendered" and "it rendered nothing" are the same result. The same trap
in the key path: a test that CALLS on_global_key passes while the window
never reaches it, because the host dispatches to its own root and the shell's
view is a child of an outer chrome. Wrap the shell's view in a parent, or the
broken arrangement is indistinguishable from the working one.
idiom_check.py asserts a patch against a claimed idiom. Its failure mode is
not "lets bad patches through" — it is rejecting correct ones, which the code's
own comment calls the worst thing a checker can do.
Because a generator can express one correct answer several ways, a too-narrow rule fails at random: a pass rate that moves between runs with no code change is a checker problem wearing a generator's clothes. Today's rate went 6 → 11 → 10 → 12 → 10 → 12 with the model untouched.
Things that have been too narrow:
transparent in
_roles.json. Two traps: the relayed module's OWN jacks are generic (a
mult's outputs are 1 2 3, role Cv, whatever is fed in), so the kind must be
checked at the cable ENTERING the chain; and relayed must be tracked
separately from reached, because when from_module is any every module is
a candidate already and "was it added by widening" is always false._port_matches falls back to labels for modules nobody
has cartographed. A port kind can list not_ports — roles that rule it out
whatever the label says. gate_out accepts SQR/PLS (an LFO's square
fires an envelope) but rules out role Audio, or an oscillator's pulse
counts as a clock at audio rate.Pitch; the vendor inference calls it Cv. A requirement naming only one
holds for Fundamental's VCO and not for ours.at_least under-states what a patch needs: a requirement names the roles at
both ends of a cable and only some appear there. patch_vocabulary.needed_modules()
derives the real list from both. The self-test cannot catch this on its own —
its fixture builds from the TOPOLOGY, inventing whatever role a requirement
names, while the model gets at_least. The instrument knows more than the
thing it tests.
Widen only with test_idioms.py watching: 54 idioms, 114 negative controls and
the textbook fixtures. The corpus has already caught a "fix" that broke
Fundamental's kick.
intent_module_plan() is a pre-model feasibility solver, so the application's
module-source setting must reach it as data. Under prefer_existing, an
installed port-complete library module outranks ForgeModular; an explicitly
named maker still wins, and the generated module remains valid when it is the
only exact jack-map fit. Do not implement "Library first" as prompt wording:
the model can ignore prose, while filtering/ranking the concrete candidate set
is deterministic and testable. Maker detection must also distinguish a
positive request from local negation (avoid Forge Modular, anything but ForgeModular) and recognize compact Rack slugs such as ForgeModular when the
user writes the display form Forge Modular.
Topology presence alone does not prove that several edges meet at the same
physical module. Use same_role_instance_groups when a technique requires one
specific mixer, delay, filter, or other role instance to participate in all
named requirements. The pre-plan must preserve repeated incoming lanes on a
single bound instance because one Rack input accepts one cable; deduplicating
two required audio_in lanes can advertise a one-input module for a two-input
sum it cannot build. Every same-instance group needs a
split_role_instance negative control, or the checker has never proved it can
reject the look-alike graph where each edge is present on a different module.
A Rack plugin is a shared MODULE, linked with -lRack and -undefined dynamic_lookup. A source list missing 25 of 28 files produces a dylib that
builds with exit 0 and 28 undefined model symbols, and Rack rejects the
plugin whole when it dlopens it — every other module with it.
So gate on the artifact, never on the build:
nm -u build-rack/rack/ForgeModular/plugin.dylib | grep model # must be empty
python3 tools/rack/test_pack_links.py # + 3 static checks
python3 tools/rack/test_rack_plugin_loads.py # does it LOAD?
That grep model is necessary and badly insufficient — see the next section
for the month it cost.
examples/forge-modular/CMakeLists.txt globs src/*.cpp with
CONFIGURE_DEPENDS for this reason. Do not replace it with a list.
The Rack SDK is developer-supplied, so most machines never configure the target
and the failure waits for whoever does have it. tools/rack/fetch_sdk.py --check reports where it is (usually ~/SDKs/Rack-SDK) — use it rather than
searching the filesystem.
-undefined dynamic_lookup is ADDITIVE, and omitting -lRack ships a plugin that cannot loadForgeModular could not load in VCV Rack for a month. Every gate was green the entire time, and the evidence was sitting in the user's own Rack log:
Could not load plugin: dlopen failed: symbol not found in flat namespace
'__ZN4rack3app10PortWidget10onDragDropERKNS_6widget6Widget13DragDropEventE'
The cause is one flag, and a misreading of the SDK. plugin.mk line 22
applies -L$(RACK_DIR) -lRack unconditionally, and line 42 then ADDS
-undefined dynamic_lookup on macOS. The flag is additive; it is not a
replacement for the link. Pulp took only the flag, so the dylib carried 140
undefined rack:: symbols and no libRack load command.
The control that settles it in one line:
for d in ~/Library/Application\ Support/Rack2/plugins-mac-arm64/*/plugin.dylib; do
otool -L "$d" | grep -q libRack && echo L || echo N; done | sort | uniq -c
# 251 L, 1 N — and the 1 was ours
Linking is not the whole fix. libRack's own install name is the bare
libRack.dylib, and a bare name never consults LC_RPATH. The load command
has to name /tmp/Rack2/libRack.dylib, the path Rack symlinks at startup —
which is what plugin.mk's dist: target does with install_name_tool -change, and why all 251 working plugins show that path.
Why five gates missed it, which is the transferable part. Every load-time
gate we had runs the plugin inside a process that has already linked libRack
— the behaviour gate links it, the patch gate links it, real Rack obviously
links it. Under flat-namespace lookup dyld resolves the plugin's symbols out
of the host image, so dlopen SUCCEEDS in all of them. The defect is
invisible in every configuration except the one the user is in.
| gate | why it was blind |
|---|---|
patch_gate.cpp dlopen | the gate binary itself links libRack |
run_behaviour_gate | links the module's .o files into a libRack-linked exe — the shipped dylib is linked a second, different way |
test_pack_links.py | reads nm -u but filters to names containing "model", discarding the whole rack:: namespace |
prove_rack_opens.py | launches real Rack, which links libRack |
test_generate_safety.py | the only test reading otool -L, and it asserted only Accelerate.framework |
So a probe for this class must do one of exactly two things: read the load
commands, or dlopen from a host that links no libRack.
test_rack_plugin_loads.py does both, builds the real pack rather than a
stand-in, and pairs every check with a control that must come out the other
way — a gate for a defect five other gates could not see has to prove on every
run that it can still see it.
On macOS the no-libRack host probe still needs dyld to resolve the plugin's
intentional absolute load command. The test must provision
/tmp/Rack2/libRack.dylib from the fetched SDK inside a scoped context, then
remove only the directory or symlink that context created. A load-command
inspection alone does not prove the artifact loads, while leaving a warm
/tmp/Rack2 behind makes later runs depend on untracked machine state.
Two producers, and both were wrong. generate.py:compile_all (the Python
tempdir build that makes the installed artifact) and
tools/cmake/PulpRack.cmake (pulp_add_rack_plugin). Linux and Windows always
linked Rack; only the macOS branch did not. Fixing one and not the other is the
standing trap here — and note the third copy: the running toolchain at
~/Library/Application Support/Forge Modular/tools/rack/, seeded from the app
bundle by install_toolchain.sh. A repo-only fix does not reach the next
real build until the toolchain is reinstalled.
What it invalidates. A patch corpus generated while the pack could not load
was never heard. 96% of the generated patches place at least one ForgeModular
module, so essentially every measurement over that corpus is about a rack with
missing modules, whatever else it also says. Check otool -L before trusting
any generation-era audio result.
A VCV Library module resolves its entitlement from a cached key at
<asset::userDir>/licenses/<plugin>.vcvkey. Nothing sets userDir for you. In
a harness that only calls dlopen it is empty, so the lookup goes to
/licenses/..., finds nothing, and the module constructs normally, runs its
DSP, and writes 0 to every output. No exception, no failed load, no log line.
That is externally indistinguishable from a module that is genuinely dead, which is the whole problem. A gate that measures outputs reports SILENT, the generation believes the module is broken, and it spends its retries replacing a part that works.
The ordering is load-bearing:
asset::userDir = <the Rack user dir> -> rack::asset::init() -> first dlopen
asset::init() called after the first dlopen does nothing. The key is read
while the plugin initialises, not when it is asked for its output.
The control that separates "unlicensed" from "dead":
ls ~/Library/Application\ Support/Rack2/licenses/*.vcvkey | wc -l
# non-zero, or the comparison below proves nothing
Then measure the same patch with and without userDir set. A verdict that
changes is a licence; a verdict that holds is a real defect. Do not
skip the first line: on a machine with no keys, both halves agree for a reason
that has nothing to do with the patch.
Every harness that loads plugins has to do this, not only the audibility gate.
test_patch.py asserts it, so a new loader that forgets is caught there rather
than by a run of patches that read as silent for a reason nobody can see.
nm cannot tell you whether a plugin links libRackThe obvious check for the missing--lRack defect above is to look at the
symbols, and it does not work. nm -m plugin.dylib | grep _ZN4rack reports
118 undefined for a correct build and a broken one alike; shipped
third-party plugins report 165. Undefined rack:: symbols are normal, because
they are exactly what the flat namespace resolves at load. Only otool -L
names the load command, and a missing load command is the defect.
patch_gate calls rack::random::init(), which seeds from the clock inside
libRack, so two identical passes over the same patch can disagree. Measured
over a corpus it is rare and real: one disagreement in 500 patch-measurements,
and always one patch flipping rather than a set.
So a replay or an A/B that reports one changed verdict has reported noise. Run the same binary twice before comparing two binaries, and state that floor beside the result. An effect has to clear it to mean anything, and the ceiling matters too: if only 13 patches in the corpus can possibly change, a result of 8 is inside a band of 2 to 13, not a rate.
@Maker and plain role words are retrieval guidance unless the user says
only (or equivalent). An exclusive maker request is a closed final-patch
constraint: allow Core I/O, reject every other maker after generation, and do
not silently substitute one on a retry. #Tag is explicit: resolve it against
the current VCV index with forgiving case/plural matching, pass its candidates
to the model, and require at least one final module for each explicit tag.
Ordinary words such as quantizer and sampler remain useful cues, not
surprise preflight refusals. Keep these checks after parsing and linting so a
model cannot claim compliance in prose while returning the wrong patch.
plugin.json, plugin.cpp, plugin.hpp, generated_modules.hpp and every
res/*.svg are EMITTED from modules/*.json. A manifest/binary mismatch is
rejected by Rack whole, so never hand-edit an output.
The emitter needs build/shape_text, built by tools/rack/build_shape_text.sh
(it needs a populated Skia checkout — headers alone compile and fail at link).
It used to default to /tmp/shape_text, and when macOS cleared /tmp the panels
could not be regenerated from a clone at all.
A module built from the app lands in the INSTALLED pack only.
tools/rack/copy_back.py lists what is stranded; --apply brings it here.
The panel emitter rejects needlessly sparse modules by default, but it receives
only the manifest—not the user's request. That once made generation prove an
exact request for a 10 HP panel and then reject the same 10 HP manifest as too
sparse. After module_intent_problems() proves that the manifest exactly
matches an explicit N HP request, generate.py records that authority in the
existing width_waiver field with a deterministic reason.
Keep the boundary narrow: an inferred or mismatched width still has to satisfy the ordinary density rule, and an existing non-empty module-authored waiver is preserved. Do not weaken or remove the emitter's generic density check to fix a request-context problem it cannot see.
OURS come from the manifest a panel was emitted from — always present, exact, never needs a scan. ANYBODY ELSE'S come from what CARTOG measured inside Rack, because a vendor's control positions exist only in compiled widget code.
Both must follow the same rule: draw what you know, skip what you don't. A
slider or a switch drawn as a circle is wrong in a way a reader cannot see, so
the manifest path maps four knob kinds to diameters and skips the rest. The
measured path used to push every control as a knob sized min(w, h), which
turned a vendor's fader into a small dial — the exact thing the other path
exists to prevent. PortMap::draws_as_knob() is the shared rule now.
An empty kind means the scan predates classification and must still draw;
refusing those empties panels that are correct today. Absent is unknown,
never knob. Same shape as the port map's scan version: "no data" and "data
saying no" are different answers and collapsing them is the bug.
CARTOG still writes lights, displays and type that nothing reads.
tools/rack/test_portmap_fields.py lists them with reasons so a dropped field
is a decision rather than an oversight, and fails on one that appears without
one.
CARTOG writes ~/Library/Application Support/Rack2/forge-portmap.json and each
scan MERGES rather than rewrites — a module measured once is carried forward
untouched forever. Entries therefore record only what the scanner of the day
knew, which is why every entry carries "scan": N (PortMap::kScanVersion).
Three states look alike and must not be confused: params recorded (known);
none, recorded by a scanner that looks for them (known — the module HAS none,
like Merge or Split); none, from one that did not (unknown). Only the scan
version separates the last two. PortMap::controls_known() is the rule; the
UNMAPPED badge reads it.
portmap_seed.entries(inv) keeps a scanned block only when its pluginVersion
equals the installed plugin's version. That strictness is deliberate: a vendor
update can add, remove or renumber ports, and a block attached to the wrong
index is worse than no block, because absent ports read as UNKNOWN while wrong
ones are trusted.
Core is the single plugin that reports no version, because it is compiled into
Rack instead of being installed as a plugin directory. An exact match can
therefore never succeed for it, and every scanned Core block was silently
dropped -- including AudioInterface2, the stereo output nearly every audible
patch terminates in, and the MIDI interfaces. The generator was wiring the one
module every patch needs while being told its ports were unknown. No rescan can
fix it either, since there is no plugin directory for measure_ranges.py --all
to visit.
So _version_admits() treats an absent installed version as a wildcard and
keeps the exact match for every plugin that reports one. Core is the only
plugin on a normal install with an empty version, so the widening is that
narrow. When a count of known ports moves by a handful for no obvious reason,
check whether a plugin's version string changed shape before suspecting the
scanner.
portmap_seed.resolved_local_path() reads FORGE_PORTMAP before falling back
to the live map CARTOG writes. A campaign that compares two port maps sets it
once; because the campaign driver shells out inheriting the environment, that
single variable pins both the inventory receipt and the generation subprocess,
so no driver argument has to be threaded through.
A set-but-unusable value raises instead of falling back. Falling back would silently mean "no local scan", which is indistinguishable from the arm the comparison is measuring against, and the run would report no difference for entirely the wrong reason.
measure_ranges.py --all copied each launch's measured map onto the live path
in place. Across the hundreds of launches a full library takes, that produced a
map truncated mid-token with a later record's bytes spliced in behind it and
doubled commas through the remainder.
The damage was invisible. portmap_seed._read returned [] for anything it
could not parse, which is correct for a machine that never scanned and very
wrong for a machine whose map is torn: known-port modules fell from about 1010
to 274 and the generator went on wiring blind with nothing anywhere reporting
it. A file that EXISTS and will not parse is a fault, not an absence.
So the copy-back goes through _install_map, which writes beside the
destination, parses what it wrote, and renames -- a rename inside one
filesystem is atomic, so a reader sees the old map or the new one and never
half of each. And _read now says on stderr when a map exists and will not
parse.
The tear originates in the scratch directory, not in the copy: Rack writes that map itself, and a launch killed part-way through the write leaves it truncated -- which is common precisely because so many of these launches SIGSEGV. An unparseable measured map is therefore a FAILED LAUNCH, retried by halving like any other, and never a reason to abandon the remaining plugins. Raising instead cost a full run at plugin 116 of 258.
Freeze the live map before any library-wide run. It is the only copy, and recovering the pre-scan state afterwards is otherwise impossible.
Random offers a prompt the user reads and edits before building, so a prompt that names something the shipped modules cannot satisfy is a promise the product breaks on its first click. Keep every entry inside what actually ships, and say what the module really does -- a sample and hold with "switchable glide" is a contract; the same entry claiming a "track mode" it does not have is not.
The pool is declared in TWO places that must agree: RANDOM_MODULE in
examples/forge-modular/app/ui/main.js and kRandomModule[] in
forge-seam/modular/modular_shell.cpp. Editing one and not the other is the
easy mistake, and the two paths are far enough apart that nothing obvious
fails. tools/rack/test_generation_eligibility.py is what catches it.
Parameter ranges (minValue / maxValue / defaultValue) come off a
ParamQuantity, which exists only on a widget Rack has built. So they can only
be measured from inside a running Rack — and for the whole life of the feature
they never were, because measuring them meant a person opening Rack and
dragging in the modules they cared about. tools/rack/measure_ranges.py
removes the person: it writes a patch, opens it headless, and reads the map
back. Two things make that work, and neither is guessable.
A headless Rack never runs a frame, so step() never fires. CARTOG's
rescan-on-count-change lives in step(), which is the per-frame callback:
Rack -h <patch> loads the patch, builds every widget, and tears the scene
down without stepping once. The widget existed and its panel loaded; nothing
ever asked it to measure. The scan therefore ALSO hangs off onAdd(), which
Rack fires as each module joins the scene, during patch load. Keep both —
onAdd is the only path a headless run has, step() is what keeps an
interactive session current.
onAdd fires per module, in patch order, so the scanner goes LAST. Verified
against Rack, not assumed: named first, CARTOG logs placed alongside 1 modules and writes a map holding none of its subjects; named last, alongside 44. getModules() already includes CARTOG itself when its onAdd runs, so
the count is subjects + 1. scan() records the count it measured, so step()
compares against what was actually seen rather than against a number a caller
remembered to set.
Every automated run must be killed, and Rack holds that against the next
one. Headless Rack has no way to exit on its own, so the harness kills it —
and the next launch decides it crashed last session and blocks in
osdialog_message → NSAlert runModal on "VCV Rack crashed during the last
session … Clear your patch and start over?". There is no window and nobody at
the keyboard, so it waits forever; the log stops dead after Creating patch manager and the run reports a library with no ranges in it. --safe is not an
escape — it skips the plugins too. Against the real user directory the failures
accumulate: measured on a working machine, 2 of 5 launches loaded the patch,
then 0 of 8.
The fix is a throwaway Rack user directory per launch (make_scratch),
with the module library symlinked in and settings.json + licenses copied
so a Pro launch stays licensed. A directory Rack has never seen has no last
session to have crashed in: 5 of 5. It also stops the scan clobbering the
user's autosave, log and open rack.
The symlink must preserve Rack's architecture-specific layout. In an
isolated proof, link ForgeModular at
plugins-mac-arm64/ForgeModular, not at the scratch user-directory root.
Rack ignores a root-level link, then a missing-module dialog can make a launch
look like a plugin proof unless the log also proves the expected module was
created. Keep a regression stub that rejects the wrong layout.
Three ways a launch dies before it reaches the patch, all of which look from the map like a library with nothing in it:
Nobody closed stdin. Headless Rack prints "Press enter to exit." and waits
on whatever terminal it inherited — forever. Every run leaves a live Rack
holding an audio device, and somebody has to force-quit it. Launch with
stdin=DEVNULL and kill the process yourself once the map is written.
Rack aborts inside its MIDI init. rack::rtmidiInit → MidiInCore →
getCoreMidiClientSingleton throws, the exception crosses a noexcept
boundary, and the process aborts about 170ms in — before any patch is
parsed and before any of our code is reachable. It is sporadic, and the
retry is the entire remedy. Two confident explanations for it have now
been measured and both are wrong, so do not reach for a third without
numbers:
| explanation | prediction | measured |
|---|---|---|
| no GUI login session | fails over SSH, fine locally | a bare MIDIClientCreate probe: 1/400 failures in a GUI session, 0/50 over SSH with no SECURITYSESSIONID |
| client exhaustion from rapid relaunch | failures cluster under hammering | 0/14 launches back to back; the single abort came from the run spaced 3s apart |
Do NOT wrap the launch in launchctl asuser. It is the obvious fix and
it is worse than nothing: it needs root, and without it the command is never
executed at all — launchctl asuser $(id -u) /bin/echo HELLO prints nothing
and exits 1 with "Could not switch to audit session … Operation not
permitted". Over SSH, 50 of 50 wrapped probes failed where 0 of 50
unwrapped ones did. It would break the one case it gets added for.
Claim this mechanism on the stack Rack prints (rtmidiInit, RtMidiDriver,
MidiInCore) and never on "the launch failed" — a diagnosis that fires for
every death stops meaning anything the moment it is right. Bound the retry
separately and tightly: each abort is a crash report on somebody's desk.
The crash prompt, above.
install_pack.sh reporting success does not mean Rack loads what you built.
It places the .vcvplugin ARCHIVE and never touches an existing UNPACKED
plugins-mac-arm64/ForgeModular/ directory — and the unpacked directory is
what Rack loads. Same version in both places is not the documented refusal
("keeping X, which is newer"); it is a clean exit 0 with the old binary still
live. Verify by hash, not by exit code, and if the directory is stale move
it aside and unpack the archive over it.
A range without a unit cannot place a number from a book. The same idea
carries four unit systems across an installed library — measured on its
filters: ALM018 runs −1..0, XFMN01 20..12000 Hz, Rain 0..1 normalised,
BattalionTone −5..5 volts. Against bounds alone, "cutoff 40 Hz" is past one
module's maximum, near another's floor, and meaningless on a third.
Scan 5 records what Rack keeps beside the bounds: unit, and the
displayBase / displayMultiplier / displayOffset that convert both ways.
displayValue = f(value) * displayMultiplier + displayOffset
f(value) = value base == 0 linear
f(value) = log_{-displayBase}(value) base < 0 logarithmic
f(value) = displayBase ** value base > 0 exponential
tools/rack/param_units.py is the conversion; use place() and pass the
unit the number came with. Converting without checking the unit is the
failure this exists to stop: 40 on a dimensionless 0..1 knob is out of range,
clamps to 1.0, and yields a filter wide open in a patch claiming to follow the
book. So a unit the control cannot read is a refusal, while a value it
understands but cannot reach is a clamp with a reason — different
situations, reported differently.
Module identity is resolved before any physical-value lookup. A name absent from the exact inventory is a model-contract error, not evidence that Cartog needs another scan. Do not measure a tag, display label, or guessed spelling; retry with the exact plugin/model pair listed by the inventory.
Written only where it is not the identity, so at scan 5 an absent
displayBase means linear and an absent unit means dimensionless; below 5
it means nobody looked. Same rule kind lives by. Emitting the identity for
every linear knob would add most of a megabyte of 0.000000 and say nothing.
Bump kScanVersion and CARTOG's emitted "scan" together — they live in
different products and forge-seam/test_seam_patch.sh greps both because
nothing else links them.
A measured range can be inverted, and a default can sit outside it. Rack's
configParam does not require min < max — a reversed knob is a legal
configuration — and a vendor may declare a default outside its own bounds (0
usually meaning "auto"). On this machine's library: 10 params across 7
modules are inverted, and 8 across 7 have an out-of-range default, about
0.1% each of 9,307 ranged params. They are faithful, not corrupt: re-measuring
reproduces them exactly. Anything consuming the map must therefore use
lo, hi = sorted((minValue, maxValue)) before normalising or sampling —
(v - min) / (max - min) divides by a negative on an inverted range and
silently mirrors the control.
A partial sweep is the outcome that looks finished and is not. An aborted
launch leaves its subjects exactly as they were, so a batch can end with some
modules still carrying an entry from an older scanner: present, parsing, and
quietly less than the rest of the map. stale_scans names them, measured
against the newest scan version anywhere in the map rather than a constant —
a constant duplicated from kScanVersion would go stale in the one place
whose job is noticing staleness.
A zero here is never evidence on its own. A wedged launch, a scanner placed
first, and a genuinely empty library all produce the same "measured nothing".
So the harness reads Rack's own log for forge: CARTOG placed alongside N and
fails loudly with the reason instead of reporting a total, and its verdict is
which models still lack ranges rather than whether a number went up. Modules
with no params at all (CV funk's blanks, our MULT) are not a shortfall.
A patch had only ever been checked as a file. It can be heard, and the way in
is not obvious: Rack -h <patch> closes stdin and exits at once, which is why
measure_ranges.py gets a scan and no audio. Hold stdin OPEN as a pipe and
Rack blocks on its "Press enter to exit." — and while it blocks, the engine
runs in real time on its fallback thread. Four seconds of stdin held is four
seconds of DSP, measured: 176,401 frames at 44.1 kHz, every time.
tools/rack/fidelity.py is the harness built on that, and
tools/rack/fidelity_probe.cpp the module it drives. Three things about the
shape are worth keeping:
The probe is a SEPARATE plugin, compiled into a scratch directory. Not a
module added to Forge Modular: a diagnostic in the shipped set would show up in
the user's browser and outlive the question. make_scratch symlinks the user's
plugins one directory at a time rather than linking the arch directory whole,
which is what leaves room to add a plugin they do not have. Nothing is
installed and the user's Rack is never written to.
Every audio interface is DELETED before the run, and the probe takes its place. The probe is patched with exactly what the interface was being fed, so what gets recorded is what a listener would have heard — and nothing in the patch can reach an output device. That is what makes the harness safe to run unattended on somebody's desk.
Ask Rack to leave; never kill it. The probe flushes its capture from the audio callback when the window fills, and from its destructor otherwise. The destructor only runs on a clean shutdown, so the harness writes a newline to stdin and waits.
tools/rack/acid_taps.py plans the synchronized acid capture in fixed order:
clock, raw pitch, post-slew pitch, accent, slide, effective external cutoff,
filter audio, final audio. It reads the acid-voice topology and the same
inventory/Cartog port roles as idiom_check.py; an unknown jack is
UNMEASURED, never "probably input zero". Feed its returned taps directly
to fidelity.instrument, whose one eight-input probe keeps every series on
the same sample clock.
There are two identities in a routed control signal and both matter. The
physical origin survives transparent multiples, attenuators, mixers, switches,
quantizers and slew modules, so pitch, accent and slide can be proven to be
three distinct sequencer output jacks. The terminal output advances through
that chain, so effective cutoff is the last external signal the filter really
receives (for example the output of a mixer combining accent and envelope),
not merely the raw accent lane. Shared physical lanes are a measured FAIL;
missing, uncartographed or ambiguous witnesses are UNMEASURED.
evaluate_capture is intentionally only a capture-contract evaluator. Its
PASS says all eight finite, equal-length series were recorded and explicitly
does not say the patch sounds like a 303. acid_behavior.evaluate is the next
layer: it derives eight-step windows from clock edges, requires two agreeing
loops and three pitches, proves accent and slide are selective, compares
post-slew motion on matched selected/unselected transitions, and uses repeated
equal-pitch accented/unaccented steps to require cutoff plus filter/final-audio
brightness or transient differences. Missing, unequal, short or loop-mismatched
series are UNMEASURED; observed silence, flat/zero cutoff, all-on selectors,
missing glide or bypassed audible response are FAIL. Structural tap-plan
failures stay separate from behavioral failures in the result. Do not weaken
this verdict into cable presence or subjective taste.
Module::paramsFromJson puts a written value through the parameter's own
bounds, so a patch that writes 0.0 into a knob whose minimum is 0.001
loads holding 0.001. Nothing reports it: the file keeps the number it was
written with, the module holds a different one, and the only symptom is the
patch not doing what it says. Found on a real generated patch —
bouncing-ball-never-bounces.vcv writes ENV Attack 0.0 — and reproduced from
both ends, by reading the engine back and by predicting it from the port map's
declared range. fidelity.will_be_clamped() names it before Rack is launched
at all, which costs nothing; the port map has to have ranges for that plugin
first (measure_ranges.py).
Four readings out of the first analyser were the instrument rather than the patch, and each one read as a finding:
Two numbers that look identical are the tell for a reading that came from the instrument. And "silent" and "unreadable" are different findings: report the recording's level, or a patch that plainly sounds reads as one that makes nothing.
~/.cache/forge-modular/ (library.json, modules.json) is one directory
for every process on the machine, and two rack harnesses started together
both refresh it. An in-place open(path, "w") truncates the file before the
new bytes land, so the second process's json.load reads a half-written index
and dies with JSONDecodeError: Expecting property name enclosed in double quotes -- a red gate that names no line in the diff and passes on re-run.
patch.py publishes both files through _publish_json(): a sibling temp file
plus os.replace, so a concurrent reader sees the previous complete file or
the new one, never a partial. Any new cache file under that directory goes
through the same helper; tools/rack/test_module_index_atomic.py reproduces
the race deterministically by probing the index mid-dump.
The single most expensive habit on this project is believing a measurement that returned nothing. Absence is what a true negative and a broken instrument look like from the outside, so a zero is the one reading that must be re-taken before it is reported. Every one of these was confident, plausible, and wrong:
| the reading | the instrument | what was true |
|---|---|---|
0 modules carry ranges | probed min/max | the keys are minValue/maxValue |
scan version: None | read a top-level field | scan is per module |
pack has no dylib | unzip | .vcvplugin is Zstandard (tar --zstd) |
the EPUB is 0 bytes | ls -s (blocks) | it is a directory; 2.3M chars of text |
no scan version in CARTOG.cpp | grepped "scan" | the quotes are escaped in C++ |
EMITS_CLOCK_TAGS is absent | grepped the wrong file | it is in patch.py |
the sweep died | pgrep -fl | ps showed it running; 8 minutes in |
no prior installers | a zsh glob | one non-matching pattern aborts the whole command |
no .component built | the same zsh glob | all four bundles existed |
render_for returns nothing | — | that one was real |
Nine of ten were the tool. The tenth — the instrument catalogue rendering zero characters — was a genuine defect, and it was only believable because the other nine had been checked and eliminated first.
The habits that catch these, in order of cost:
find reported SDK directories empty while ls -A showed contents,
because find does not follow symlinks. pgrep said the sweep was dead
while ps said it was running. The disagreement is the information.setopt NULL_GLOB or use find. zsh aborts an entire command when any
glob matches nothing, so a single stray pattern silently discards the output
of everything beside it. This happened four times in one session.A run log said EnvelopeArray could not receive a gate. Ten minutes later the
live inventory said it could. The right question was "what landed in
between"; the question actually asked was "was my evidence stale". A fix had
been committed between the two, so a real defect was declared imaginary and
another agent was told to stop working on it.
git log --oneline -5 would have settled it in one command, and on a tree
several agents are committing to it is never a wasted one.
Both readings were true of their moment. A log records what happened once; a query reports what is true now. Neither is a statement about the other, and reconciling them means finding the commit, not picking the more recent instrument.
The corpus audit asked, for every jack a real cable lands on, whether any port kind would accept it — and reported 0 failures in 319 cables, which read as "the matcher is healthy".
It is not a health measurement. cv_in accepts nearly anything: a jack named
Wobble classifies as cv_in quite happily. So the question could barely fail,
and the zero it produced said almost nothing.
The defect being hunted was never "no kind accepts this" — it was "the WRONG
kind accepts this". Gate 1 CV was not unplaceable; it was placed as Cv,
confidently, by a vocabulary that checks Cv first. A test that asks whether
something matched cannot see a thing that matched incorrectly.
When designing a detector, ask what a defect would look like in its output before running it. If a defect and a clean result produce the same reading, the detector measures nothing however impressive the sample size.
~/Library/Application Support/Forge Modular/corpus/patches/
provider.json the acquisition descriptor: which third party, its query
API, the delay its robots.txt asks for, contact string
index.json per patch: id, title, author, url, licence, licence slug,
sha256, tags, categories, fetched_at
patches/*.vcv the bodies — Zstandard, `tar --zstd -xf … patch.json`
The source is named in provider.json, not in the repository. Which third
party the corpus came from, and the terms that governed the fetch, are
acquisition provenance and live with the evidence. FORGE_CORPUS_ROOT
overrides the location; tools/rack/patch_corpus.py refuses to fetch without a
descriptor and needs none to read what is already held.
Outside the repo entirely, and deliberately so — ditto copies a directory
rather than git's view of it, which is how 113 MB of third-party reference
material once reached a signed installer.
Licence never opens the public-export boundary. Third-party PDFs, images,
patch bodies, titles, uploader identities, URLs, ids, hashes, source prose, and
full per-patch module/cable topology remain private regardless of their
licence. tools/rack/corpus_export.py is not a backup command: it
writes one fixed-schema report containing only source-neutral aggregate counts,
abstract functional-role vocabulary, independently worded criteria, and
corroborated generic signal-prior buckets. It refuses an existing destination
so an older source-shaped export cannot survive beside a new public report.
Private backup or migration belongs in private-repository tooling; the research
tools continue to read the machine-local corpus directly. Keep
rack-corpus-export registered and its malicious synthetic fixtures green when
changing this boundary.
Fetched via a public, documented query API under the terms recorded in
provider.json. robots.txt governs: the delay it asks for is honoured to the
second and a Disallow means no fetch at all. Every patch carries its own
licence in the API response, which is a stronger permission signal than a
site-wide document because it comes from the rights holder. The licence is
recorded per patch, so anything non-permissive can be excluded from storage
later without re-fetching.
What it was predicted to find: port-matching defects in bulk, over real usage. What it found: nothing, for the reason in the section above. What it was actually worth: it exposed a wrong diagnosis. Building the audit is what surfaced that a "defect" had been read off a superseded log — which stopped a permanent loosening of the gate that would have been made for a reason that never existed.
That is worth recording honestly rather than as a success. The corpus is more
promising as usage-inference for unmapped modules — if many patches wire a
module's out0 into a V/Oct input, that is a pitch output, and that is
knowledge about the ~40% of the library CARTOG cannot reach because nobody has
it installed. It is a prior, never a measurement, and must never override a
real scan.
usage-priors admits a port hint only when several distinct contributors
independently wire a jack the same way, so one contributor cannot corroborate
themselves by uploading twice. That makes the contributor identity, not the
floor, the part that decides whether the lane can ever admit anything.
An identity that reads a field the stored evidence does not carry does not
error. Every observation falls into the same anonymous bucket, support pins at
exactly 1, and any floor above 1 becomes unreachable. The lane then prints
corroborated priors: 0 with a four-figure quarantine, which reads as a
finding about the corpus rather than a broken instrument, and the honest-looking
number invites the wrong fix: lowering the floor. The floor was never the
problem.
So the lane now reports degenerate_support and prints a warning when every
observation has support exactly 1, because a corpus that genuinely lacks
corroboration still shows a spread of support counts. When you change anything
about how contributors are told apart, re-run the sensitivity control rather
than reading the headline: run it once each at --min-support 1, 2 and 3 and compare. A
real corpus decays (many at 1, fewer at 2, fewer at 3); a broken identity falls
off a cliff (everything at 1, nothing above it).
The identity is a comparison token, never a label: contributor names are hashed before they are counted, so nothing downstream can print or persist one even by accident. Tests for this must be shaped like the stored records rather than like the writer's output, since a fixture that supplies whichever field the code happens to read agrees with it by construction and proves nothing.
corpus_audit.py prior-accuracy scores the inference where the answer is
already known. The lane infers an unmapped port's class from the mapped end of
a cable, and by construction nothing can check that: the port is unmapped
because there is no scan of it. A cable with both ends mapped is the
held-out case, so run the same inference there and compare against the port's
own class.
Measured over the current corpus:
| min-support | scored | correct | wrong |
|---|---|---|---|
| 1 | 400 | 324 (81.0%) | 76 |
| 2 | 128 | 117 (91.4%) | 11 |
| 3 | 60 | 59 (98.3%) | 1 |
| 5 | 16 | 16 (100%) | 0 |
Two things follow. The admission floor of 3 is not arbitrary: it is where a
four-in-five guess becomes a fifty-to-one one. And accuracy climbing with
support is the control on the support count itself - if contributors were
not really being told apart, support would be noise and the line would be flat.
prior-accuracy warns when it stops climbing, for exactly that reason.
Read the number as an upper bound and say so when quoting it. A port has a known class here only because it is mapped, and mapped ports are plausibly better named than the unmapped ports the lane actually serves. This scores the inference, never the module: a prior remains a hint that a real scan overrides.
Say this plainly before quoting the accuracy table at anyone: the report is proposal-only and no code reads it. Count the readers rather than assuming one exists. The lane produces good hints that currently go nowhere, so the remaining work is a consumer, not more inference.
When that consumer is built, it does not belong in the port map, and the port
map's own docstring says why. portmap_merge.hpp folds a fresh scan in by
replacing a re-measured module's whole block, so anything inferred that is
stored there is erased the moment somebody scans the module. Ranges belong
there because they ARE the scanner's fact; a classification is not, which is
exactly why affordances are kept in their own cache instead. A prior is
inferred, so it follows the affordance precedent, not the range one.
The precedence a consumer has to preserve, strongest first: this machine's own scan, then the shipped seed, then a prior. A prior may only answer where the port map has no answer at all. It never fills a gap inside an entry the scanner already replaced, because "the scanner looked and found nothing" and "nobody has looked" are different answers and merging them is how an inference comes to outrank a measurement.
The source shelf is large enough that opening books one by one is no longer a credible research workflow. Build its machine-local SQLite FTS5 index with:
python3 tools/rack/corpus.py --index
python3 tools/rack/corpus.py --query "sample and hold evolving random voltage"
# Structured results for an agent:
python3 tools/rack/source_index.py query "acid accent slide" --json
source_index.py indexes only .md, .txt, and .sc documents already under
~/Library/Application Support/Forge Modular/.corpus (or
PULP_RACK_CORPUS). It preserves a stable source key, work identity,
extraction identity, page/section locator, source hash, and passage hash.
Re-running the build hashes every source, updates changed documents, removes
vanished ones, and leaves unchanged rows alone. The SQLite schema is versioned
and fails closed if a newer tool has already migrated the database.
Recovered subset-font text that labels itself machine output, imperfect is
kept in the corpus but gets no searchable passages. Its fragments can look
like synthesis vocabulary and manufacture excellent-looking false hits.
source_index.py status lists every such zero-passage source explicitly. That
means "we possess this source but need page-image review/OCR," not "the work has
nothing useful in it." Allen Strange currently falls into this category; use
its pinned page images and candidate-evidence workflow until recovery is clean.
The database lives inside that external corpus, never under tools/rack.
It contains the copyrighted source text needed by FTS and therefore follows
the same non-shipping rule as the books; it is forced to user-only file
permissions on every open. A query returns bounded snippets plus locators and
hashes; it has no dump command.
Most important: a search hit is evidence to inspect, not guidance to use.
Retrieval may suggest a candidate technique, repair, negative control, or
parameter relationship. It does not bypass knowledge.py's candidate
quarantine, canonical-claim dedupe, source anchor, or guided validation. The
generator still receives only admitted records. This boundary is why the
index can be broad without turning unreviewed prose into prompt cargo.
_port_matches compares a jack's label to a role's label list whole:
label.upper() in ok_labels. Vendors name jacks descriptively, so this fails
on every jack whose name says more than the bare word:
"PITCH CV (1V/OCT)" vs cv_out labels ["CV", ...] -> no match
"GATE 1 CV" vs gate_in labels ["GATE", ...] -> no match
"SQUARE" vs clock_out labels ["SQR", ...] -> no match
Each was found, diagnosed and fixed separately, as though they were unrelated bugs, and each fix was a new entry in a list. The fifth was found only because a run log printed the rejected jacks beside the complaint:
this patch's envelope CANNOT receive it, however it is wired:
CVfunk/EnvelopeArray (its inputs are ..., Gate 1 CV, Gate 2 CV, ...)
When a fix is an entry in a list, ask what the list is compensating for.
A per-label fix costs an hour and buys one module; changing the comparison
costs the same hour and buys the class. The reason to be careful rather than
bold is real — a wider match is how a gate becomes a rubber stamp, and
clock_out's not_ports: ["Audio"] exists because an oscillator's audio-rate
pulse once read as a clock. But "careful" means measure the blast radius and
sweep the corpus, not "add one more label".
Every shared format here had more parsers than anybody was checking, and the gap is invisible because each reader works perfectly in isolation:
| format | readers | tested before |
|---|---|---|
the success line built N modules, M cables → path | 4 | 1 |
the generators' endings (raise SystemExit) | 4 | 0 |
| the port map | 3 | 1 |
the .why.json sidecar | 1 | 1 |
The success line is parsed by drive_app's regex, the app's scan for a .vcv,
and a sed of its own in each prove script — so pinning one is pinning
nothing, and a prove script that finds no path reports a PASS with nothing to
open. The port map is read by the app for drawing AND by patch.py for real
jack names, so a renamed field costs every explanation its labels silently.
So before trusting a format: grep -rl the filename or a distinctive phrase,
count what comes back, and check them all against the producer's own source.
Never against a fixture — a fixture written by the same person as the parser
agrees with it by construction, which is how three of these stayed hidden.
Four things watch the generator's output and turn it into a verdict: the app's
BuildMonitor, drive_app.py, prove_surfaces.sh and prove_idioms.sh. All
four were wrong in the same way, and the symptoms differ enough that they look
like four unrelated bugs:
running and the app watched a dead build
forever — no verdict, no artifact, nothing to open.tail -2 | head -1: an arbitrary line. That
produced a mid-sentence fragment for a login failure and a bare ^^^^^^ for
a crash, and both times somebody opened the log by hand.Three questions are worth asking of any of them:
raise SystemExit in patch.py and
generate.py is an ending.tools/rack/test_generator_endings.py asks all three across the generators,
the monitor and drive_app. But a textual match is only a screen: it once
reported an ending as "recognised" that classify() returned as progress,
because it had matched a SUCCESS rule — which is worse than matching nothing,
since a failed run would then report done and offer an artifact nobody wrote.
The authoritative check runs classify() itself.
reason.sh is the shared "why did it stop" shim, for the same reason cap.sh
is shared: this rule already existed twice and both copies were wrong.
input N changes no output when connected versus muted is the gate telling the
truth about a module whose DSP is correct. The commonest cause is a modulated
param whose default sits in a flat region of its law, where the CV is added
to a knob position that cannot move.
AnalogVcfT's minimoog voicing is the one that bites:
kMinimoogResonanceKnots { 0.0, 0.25, 0.50, 0.60, 0.75, 0.90, 1.0 }
kMinimoogResonanceValues { 0.079, 0.079, 0.079, 0.52, 0.90, 0.97, 1.0 }
Knob 0.0–0.50 all map to 0.079, so a resonance CV on a knob defaulted at 0.30 produces bit-identical output — measured at every tone from 110 Hz to 8 kHz, not merely "too small to detect". It is a measured calibration curve: a real Minimoog's resonance control does nothing over its lower half. Minimoog only — Juno, Jupiter and Prophet are monotonic from zero and need no floor.
Measured against the gate's 1e-6 threshold with the gate's own +0.05 CV:
| knob default | rms_diff | |
|---|---|---|
| 0.30 | 0.000000000 | fails |
| 0.50 | 0.126 | passes |
| 0.55 | 0.100 | passes |
| 0.60 | 0.078 | passes |
The param sweep cannot catch this. Section 1 of the gate moves each knob across its FULL range, where such a law does move, so the knob reads as live. Only the default position is dead.
So the gate re-probes with params pushed to 0.75 of range and says which it is:
a wired jack whose default is at fault, or a genuinely dead one. That diagnosis
is retained for an explicit --retries 1 follow-up. The default remains zero:
pressing Build authorizes one model call, not a second hidden call, even when
the rejection is actionable.
Proving a fix here needs the real gate, not a filter probe. Driving
AnalogVcfT directly proves the FILTER responds; it says nothing about whether
the gate accepts a module built that way. Use the replay harness, which runs
validation, compile, behaviour and packaging with zero model calls:
python3 tools/rack/generate.py --response-file <saved-response> "<prompt>"
Saved responses live at $TMPDIR/forge-module-attempts-*/attempt01-model-response.txt.
Change one line in a copy, replay both, and the control/test differ by exactly
that line.
The role-aware probe must also make separate CV jacks independently observable.
Do not drive every Cv input with the same 0..5 V square: a signal edge and its
rise/fall modulation then move together and can pin an exponential time law at
an endpoint, falsely reporting a working knob or jack as inert. Keep generic CV
stimuli bipolar, centred on zero, and phase-separated by input index; do not
assume input zero is the signal. test_inert_input_cause.py includes a real-gate
slew fixture for this exact failure class. When it trips on a paid response,
retain and replay that response before changing the prompt or calling a model.
Generated source is normalized before compilation. When it uses a class from
the curated module-DSP registry without an owning public header in the
normalizer-owned prefix, toolchain_headers.py prepends that header and retains
the normalized source. Replay the saved response before spending a repair call;
a missing known include is deterministic toolchain work, not a model question.
The inert-control gate uses two probes deliberately. The driven probe connects
role-aware inputs, while the normalled/internal probe sets every input's channel
count to zero. Supplying quiet zero volts while leaving channels=1 is still a
connected Rack cable and bypasses internal clocks and other normalled sources.
Keep test_clock_behaviour_gate.py's normalled-clock control green whenever the
probe wiring changes.
The module gate is a consumer of Pulp's public signal headers, so it must also
inherit the platform link requirements exported by pulp::signal. On macOS
that includes Accelerate for vDSP-backed FFT users. A gate compile or link
failure means the harness is unavailable; stop immediately, retain its exact
diagnostic, and spend no model retry. Only a successfully built gate can reject
generated DSP and provide actionable retry context.
Generation writes into the pack, so never run it inside a checkout you are
using as FORGE_MODULAR_TOOLCHAIN_ROOT — CMake refuses a dirty toolchain and
the next Forge configure fails.
The patch-side twin of the inert-input case above, and it is worse, because nothing in the patch distinguishes the broken version from the working one.
Measured, not reasoned: the same edge into Fundamental/VCF in:0 reads
CAUSAL at 1206x the detection floor with the depth control open, and
NOT_CAUSAL at 0.00x with it at its default of zero. Same cable, same
analyzer, structurally identical patches. A label reader calls both a
modulation edge, and every structural check in lint() passes on both. So a
patch that cables a CV input without opening its depth control looks
correct, lints clean, and does nothing — the person who asked for it gets no
error, just a result that is inexplicably boring.
Module authors default depth controls to zero because that is the safe
default, so this is ordinary rather than exotic. Measured over 184 generated
patches (69 prompts), 21.7% of prompt families shipped at least one dead
edge — and they were the modulation prompts: slow vibrato on a sustained
note built LFO → attenuator (set to a careful 0.06) → VCO Frequency modulation, and left FM amount at 0. There is no vibrato in that patch.
tools/rack/cv_depth.py handles both halves, and only one of them needs
evidence:
open_depth_controls is the fix, and it runs inside prepare_and_lint
before anything judges the patch. Cable a CV input, get a nonzero depth.
Nonzero by default, never nonzero by force — a value the model or the
prompt authored is left exactly as written, including a deliberate zero.
Opening a control is benign, so a nominated association is good enough to
act on.dead_edge_errors is the report, and it can reject a patch, so it only
does that where a witness receipt rendered audio
(cv-depth-receipts.json). Everything else reaches the reader through
explain() and fails nothing.Three verdicts, never collapsed. FAIL needs a witness that graded
blocking. ADVISORY is a name heuristic — its precision beyond the ground
truth it was built from is UNMEASURED, which is exactly why it may not
reject a user's patch. UNEVALUATED means the module has depth controls and
we do not know which one governs this input; failing a patch for a gap in our
own coverage would punish the user for our missing measurement.
Nominated.blocking returns a literal False and the class carries no field
that reaches it. That is deliberate: a guess must be structurally unable
to reject a patch, not merely policy-bound not to. test_cv_depth.py asserts
it at confidences past anything the rules emit.
Every association rule abstains rather than guess. Two candidate controls
for one input is an abstention, not a coin flip — that is how a heuristic
acquires confident wrong answers. The consequence is real coverage loss:
Cutoff CV (1V/oct) and Frequency modulation defeat word matching, so a
sole-candidate rule pairs by counting when a module offers exactly one
modulation input and exactly one depth control. It reaches the two commonest
dead edges in the corpus and is still NOMINATED.
What this does NOT reach. The rule only fills a control the patch never
mentioned. Where the model explicitly writes the depth param to 0.0 — which
is 8 of the 9 dead edges surviving the fix — the rule correctly stands down
and the edge is only reported. That residue is a prompt-contract problem, not
a default-value one.
Coverage is the honest caveat. An association is known for only ~9% of cable-fed inputs, so 21.7% is a floor, not an estimate. Before quoting any dead-edge rate, print how many edges were evaluable: a low number here is usually the instrument, not the world.
The sibling of the dead-edge rule, and the one with evidence at scale. Measured as a census over every silent corpus patch carrying a stopped module: the unmodified control held silent 62/70, flipping the run flag to true woke 22, and a sham arm flipping the same number of blob booleans on keys that are not run-shaped woke 1/70 — a 25x contrast. The modules that woke are dominated by master clocks, so one flag silences the whole patch rather than one branch.
The state lives in each module's saved data blob, which decodes as ordinary
JSON with self-describing keys. running is the commonest.
An allowlist writes; a heuristic only reports. Name shape was the right instrument for measuring a corpus of other people's patches, where the module set is unbounded. It is the wrong one for a writer: the generator knows exactly which modules it emits, so it can be exact. Two off-valued keys, both found by reading a family's key set rather than by pattern, show why:
ImpromptuModular/clockMaster: false marks the clocks that are NOT the
master. Setting several true creates a conflict — and this is the same
plugin whose clocks dominate the wake list, so it is inside the blast radius
rather than at the edge of it.Valley/frozen: false is a reverb's healthy state. True silences the
patch. The same shape as the bypass inversion.From a name, running on an unread family is indistinguishable from
clockMaster. So TRANSPORT_RUN_KEYS (plus any run field the curated registry
declares) is what may be written; a run-shaped key anywhere else is
reported as SUSPECTED and left alone. Extending the allowlist is a short read,
not a guess: key counts look prohibitive only because one concept is indexed per
track or step (manualBeat-0-3, id_t3_fadeRate), and collapsing the index
takes 364 keys to 91 and 692 to 100 — Impromptu, CountModula and Valley are 30,
20 and 8.
Three sign errors, each of which INVERTS a result rather than weakening it:
bypass and muted are not run flags. bypass: false is the healthy
state; folding it in makes a writer switch every healthy module into bypass.
Note where the exclusion actually bites: the run-word set is narrow enough
that a bare bypass never qualifies anyway, so NOT_RUN_WORDS earns its
keep only on compound keys like bypassRunning. The same argument keeps
enabled, active and on out — a false there is routinely the legitimate
default of an optional feature. clockMaster and frozen are denied by
NAME as well, because the word set happening to contain neither CLOCK nor
FROZEN today is an accident: widen it and the exclusion would vanish
silently. Proven rather than asserted — widening RUN_WORDS to include
CLOCK/MASTER/FROZEN leaves the suite green, and only dropping the by-name
denial as well turns it red.[true] × 24, or a run: [1],
must not read as stopped. Held in two independent places, so breaking either
one alone leaves the behaviour correct.0/1 may
ignore a JSON boolean — a silent no-op presenting as "starting it does
nothing", which is a false null rather than a visible failure.Why this is a writer and not an audit, and why that reading is the OPPOSITE
of the dead-edge rule's. cv_depth treats an authored zero as intent and
stands down. run_state treats an authored false as the defect and flips it.
The difference is whose patch it is: a person who stops a clock and saves meant
to, and their saved patch is evidence of a choice; a generator has no such
habit and no way to notice it emitted one. The single channel that can ask
for a stopped module is module-state-overrides.json, and only a value
declared stopped there is honoured.
The registry does not close this on its own — verified, not assumed.
materialize_module_state applies data_defaults with setdefault, so it
fills an ABSENT key and never displaces one the model already wrote:
AS/SEQ16 registry declares running: true
authored {"running": false} → after materialize_module_state → false (!)
authored key absent → after materialize_module_state → true (control)
So the one entry that exists to force that sequencer running is defeated by the
model writing the field itself. run_state.start_stopped_modules runs after
materialization and closes it.
It never invents a run key. Where a module carries no run-shaped key we do not know the key's name, its type, or whether the module has one — and writing a guessed key into a saved-state blob is exactly the silent no-op of sign error 3. That case is UNEVALUATED.
And UNEVALUATED is nearly all of it. Across 183 generated patches: 107
clock/sequencer instances, of which only 5 carry any run-shaped key. The
other 102 are decided by the module's constructor default, which no patch-side
analysis can see. The generator emits almost no saved state at all — three
distinct data keys across 1350 modules — so this class is nearly not
expressible in its output. The rule is a guard against a defect the model can
introduce and against the registry gap above, not a fix for something currently
happening. Before quoting a rate here, print how many modules were evaluable.
That test was written, was correct, and exited 1 — and nothing ever ran it. It carried no ctest registration, so it drifted by 24 unmatched endings while reporting them accurately to an empty room. The app hung on the plainest possible case: the curation gates end a run before anything reaches the model, so a mistyped prompt produced a spinner that never resolved, with the reason sitting in the log.
The list and the checker are therefore one change, never two. Fixing the
endings without registering the check just resets the clock to the next drift.
It now runs as ctest rack-generator-endings (labels rack;contract, ~0.04 s,
pure Python — no Rack SDK, so it must never be gated on PULP_HAS_RACK).
Two things worth knowing before trusting a green run of it:
build_monitor.cpp and
confirm the ctest fails, then restore. tools/scripts/confirm_failure.sh
automates that loop and handles the same-second-mtime trap that makes a
hand-rolled version report a false verdict in either direction.refusal, not an error. Nothing broke:
the request was understood and declined, and the wording is the whole answer
because no model call was made. outcome_of ranks refused and failed alike
as terminal, so either ends the spinner — but only one of them tells the
truth on screen.Registration is the general point, not a detail of this test: 42 of the 48
tools/rack/test_*.py harnesses have no ctest entry. Some gate on Rack or the
network deliberately; a pure-source checker never should.
tools/rack/with_app_model_selection.py -- <command> snapshots Forge Modular's
saved engineering provider, model and reasoning effort into the same FORGE_*
environment the app gives generate.py and patch.py. Use --settings <path>
for an isolated settings file. It clears the unselected provider's model and
effort variables before replacing itself with the command, so switching from
Claude to Codex (or back) cannot inherit stale values.
This wrapper deliberately has no built-in model or effort default. It follows
only saved role-to-default fallbacks, migrates the legacy gpt-5.6 display
alias to gpt-5.6-sol, and refuses a missing, corrupt, incomplete or unsupported
selection before the command starts. A headless benchmark that silently picks a
different agent than the app is not reproducible evidence.
Codex generation must also be a sealed execution: pass --disable apps,
--disable plugins, and --disable skill_search before exec. Those
interactive catalogues consume the bounded skills context and can reject a
generation before its contract reaches the selected model. Keep the exact
argv covered by both test_generate_model.py and test_patch.py.
The monitor learned to classify every ending and the shell still reported one
line of it: BuildOutcome::failed narrated headline(), which by construction
returns the ending line and nothing after it. When the generator gives up it
says considerably more — what was asked for, what the patch it is handing over
does not meet, where that patch is, where the attempts behind it went — and all
of it was discarded a second time, one layer above the bug that discarded the
patches.
BuildMonitor::closing_block() returns the ending line and everything
after it. It scans BACKWARDS for the last ending, not forwards for the
first: a model call that failed and was retried is an ending-shaped line in
the middle of a healthy run, and starting there hands back most of the
transcript as though it were the verdict.failed branch also opens the handed-over patch (open_patch_file +
save_project_for, as done does). Without it the patch is on disk, the
skeleton is still animating, and nothing on screen can reach it.~/Library/Application Support/Forge Modular/runs/<stamp>.log, and the only
way to read a failure was to already know that path. The failed branch says
it.format_failure_report() is a free function taking a RunFailure, so a test
can build the copyable block with no shell, no chrome and no clipboard. It
reads the log back from disk rather than rebuilding it from the lines the
monitor kept: the monitor drops blank lines and holds only what it polled, and
a report that quietly differs from the file it names wastes an hour when
somebody compares the two.prove_surfaces.sh runs the CLI, the app by clicking Build, and a DAW.
FORGE_HOST=<host> re-execs over SSH. It refuses rather than falling
back when the toolchain is missing there — it silently ran locally and
printed PASS for a long time, so any older "proved on the M5" was a claim
about the machine it ran on.The signed app's qualification inputs use a deliberately non-duplicated helper
layout: FORGE_BUILD_INFO and build/rack_patch_decode are under Resources,
while the executing Python toolchain is Resources/tools/rack. Qualification
may treat that decoder as the executing candidate only when the build-info file
has its exact name and the paths resolve beneath a real
*.app/Contents/Resources boundary. An installed or developer toolchain outside
that exact bundle layout must still carry its own rack_patch_decode,
identical to the candidate by the canonical content identity, so a missing
helper cannot pass merely because some bundle was named. Decoder identity comes
from the bundled
examples/forge-modular/binary_identity.py, not a raw file digest: Developer
ID signing replaces the Mach-O signature after the stamp is authored. Raw bytes
may differ only in that excluded signature; executable-content drift must still
fail qualification.
forge-seam/modular/ holds the shell and its views; forge-seam/patches/
holds the chrome diff and the commit it applies to.
git -C ../forge worktree add /tmp/forge-cur "$(cat forge-seam/patches/BASE)"
forge-seam/populate.sh /tmp/forge-cur # seam -> worktree, and verifies
forge-seam/sync.sh # worktree -> seam
Both derive their file lists — populate from what the seam carries, sync from
the foreach(_forge_modular_src) block in the worktree's CMakeLists. They used
to carry hand-written lists that drifted, and CMake skips a missing source with
if(EXISTS) rather than failing, so a file populate missed was a model that
quietly was not there.
Run sync.sh before finishing any session that touched the shell. The
worktree lives in /tmp and macOS clears it; a test written there and not synced
is simply gone.
A mention has to be able to span spaces, because a maker's name usually has one
in it. Both halves of this originally capped the span with a literal — the @
list allowed two spaces, and brand_mentions() in patch.py tried phrases of
three words, then two, then one — and both literals are wrong for the same six
of the 375 makers in the real library: Studio Six Plus One, The All
Electric Smart Grid, Path Set x Omri Cohen, Mathematics and Music Lab
(MML), Jasmine & Olive Trees, Autodafe - REDs FREE. Naming any of them
resolved to nothing at all, in silence, from either surface.
The two sides fix it differently on purpose, and neither adds a bigger number:
patch.py reads the span from the library. brand_phrase_span() is the
longest maker name in the directory, in words, so the scan is exactly as wide
as the data it is scanning for and stays correct when a maker with a longer
name publishes.@ list has no cap at all. Only a newline ends the token. What ends
a mention is that it stopped matching, which the caller already asks and
which is strictly better evidence than a word count: right for
@The All Electric Smart Grid, and right for @vco into a filter where a
word count is only right by luck.Two more makers were unreachable for a different reason, and both surfaces had their own version of it:
patch.py dropped every token containing a slash, because
@CVfunk/Sphinx names a module. Catro/Blanco (8 modules) and p.s.F/X
(7) have one in their names, so neither could be named at all. The exemption
is checked against the brand directory, not guessed: a slashed token is a
module unless it is a maker's name.fold() in module_catalog.cpp removed a space, a hyphen and an
underscore and nothing else, so those same two could be reached only by
typing their punctuation exactly. It now drops every ASCII non-alphanumeric,
which is the rule fold_name was already applying on the other side.
Keep the non-ASCII code points: std::isalnum in the C locale answers no
to a UTF-8 continuation byte, so a byte-wise test shortens Instruō to
instru and makes it collide with anything else starting that way.The two folds are one behaviour, and "ASCII case-folding" is not enough to
make them so. fold_name is str.lower() plus str.isalnum(), both
Unicode-aware. fold case-folded only ASCII, so ÄSK — a real maker — folded
to äsk in a prompt and Äsk in the list, and @äsk found nothing. The
other three non-ASCII makers (Instruō, Hügelton Instruments, nozoïd)
agreed only by luck: all three are already lowercase, the one quadrant where
the two rules cannot differ. A sweep that types every name the way its maker
writes it will never see this — vary the case.
fold now walks code points, not bytes (half of Ä is not a character) and
case-folds Latin-1 Supplement, Latin Extended-A, Greek and Cyrillic. Every range
was checked against Python's str.lower() code point by code point; the
exclusions are the whole of the disagreement (U+00D7 ×, U+03A2 unassigned,
U+0130 whose lowercase is two code points). Outside those ranges a code point
passes through in the case it was written, and the guard for that is on the
corpus, not the code: check_maker_names_as_written sweeps the real index for
an uppercase character outside the covered scripts, or any non-alphanumeric
non-ASCII one (which Python drops and fold keeps). Widening the C++ side to
full Unicode would mean a category table no maker needs yet.
It is case-folding, not transliteration. Instruo and Instruō must stay
two makers; merging them is worse than the bug being fixed because it is silent
and it changes what an already-stored token resolves to. Assert that as a
count of maker rows, never as which row leads — with the diacritic folded
away both names match exactly and the winner is whichever the catalogue lists
first, which an assertion about hits[0] passes straight through. That version
was written, mutated, and did not fail.
Pinned identically on both sides: SEAM_FOLDS in test_patch.py and the table
in test_chrome_no_leak.cpp. Two implementations in two languages can only be
held together by the same literals on both, never by a third implementation that
could itself be wrong.
Sweep the real index before believing a matcher change is complete. Every one of these was found by asking whether all 375 makers can be named, not by reading the code:
brands = sorted({v["brand"] for v in cat.values()})
[b for b in brands if b not in P.brand_mentions(f"a patch with @{b} in it", cat)]
What pays for the uncapped span is that matching is monotone — a query matches when its folded form is a substring of an entry's folded name, slug or maker, and every longer query contains the shorter one, so a token that matches nothing cannot grow into one that matches something. The overlay remembers the token it abandoned; anything that token grows into is dismissed with one string compare instead of a search. That is worth having: measured over a library the size of the real one, a 500-character dead token costs 24 ms a keystroke to re-answer (2 chars: 1.1 ms, 120 chars: 6.6 ms), so deleting the cache is not an option and neither is capping the span by guessing.
Rely on that property if you touch matches(): widening it to a fuzzy or
subsequence match would break the proof and put an unbounded search on every
keystroke.
And the proof is about a FIXED corpus, which this one is not. all()
reloads itself when the index changes, by design, because on a machine that has
never had an index the file arrives about a minute after launch. A cached
"nothing matched" therefore expires: somebody types @Catro Blanco into an
empty library, the index lands carrying that very maker, and a guard with no
invalidation keeps the list shut against it — silently, and recoverable only by
deleting back past the abandonment, which nobody will guess. abandoned_ is
paired with catalog_generation(), which all() itself bumps on reload;
deriving that number from a second stat of the same file would let the counter
and the list disagree about when the change happened.
The general shape, worth carrying past this file: a cache of a negative answer needs the same invalidation as a cache of a positive one, and an optimisation justified by a proof is only as durable as the proof's unstated assumptions. Write the assumption into the comment so the next reader can see what would break it.
@ list ranks two kinds of row on one scale, and installedness is not one of themModule rows and maker rows are merged into a single list, so one ordering
decides both. The tiers are named in module_catalog.cpp (kExactName,
kAlias, kMakerNamed, kMakerPartial, …) with static_asserts, because as
bare integers "a partial maker name sits under the alias tier" was a
relationship between a 3 in one function and a 2 in another that only
somebody reading both would see.
The trap is the front tier, not the ranks. It was "installed, an exact name,
or a maker", and a maker row is always in it — so on a machine without
Audible Instruments, @br put any maker starting with those two letters above
Braids, whose alias match is a whole tier stronger. Being uninstalled is
precisely the state a mention exists to remedy, so it cannot be the thing that
decides which row leads. The front tier is now is_identity_match — the whole
name typed out, or the alias — plus installed rows and maker rows.
The old test for this rule installed Audible Instruments first, which is why the
hole stayed open. If a mention test flips installed = true on the fixture to
make the assertion pass, that is the bug, not the setup.
Widen that tier only against the real index, never by reasoning. Adding the
name-PREFIX tier is the obvious next step and it is wrong: measured over the
4,735-module index, @br then answers with thirty modules that merely begin
with those letters (Breakout, Branes, Broadcast) and Braids and every maker
row fall out of the six visible rows. A prefix is too common to lead. Both
directions are pinned by tests, so the decision fails loudly rather than drifts.
The front tier deliberately overrides rank — an installed kMakerPrefix row
leads an uninstalled kNamePrefix one — because it answers "what may lead the
list", not "what matched best". That is by design and predates this work.
The visible consequence, worth knowing before somebody reports it as a bug: an
alias-tier module now leads a partial maker row, so at @CV the modules
whose slugs begin CV (Core/CV-Gate, shown as "Gate to MIDI";
Autinn/CVConverter, shown as "Conv") sit above the "CV funk, 43 modules" row.
That is kMakerPartial sitting under kAlias, which is what the tiers have
always said; it was simply never true for an uninstalled module before.
Projects store the prompt (prompts in project.json), so whatever the @
list inserts has to go on meaning the same maker against a library index rebuilt
since. It does, and only because what is inserted is the maker's own display
name and nothing else: no plugin slug, no id, no module count, no syntax we
invented. brand_mentions() resolves it again from scratch each time, so the
same token picks up plugins the maker has added since and drops ones that are
gone. Anything that embeds a snapshot of the library in the token — the count,
the slug it happened to come from — silently stops resolving the first time the
index is rebuilt.
AppKit offers every key-down to performKeyEquivalent: before keyDown:, and
that path asks the focused view before the root's on_global_key. The
composer is a multi-line TextEditor, so it claims Up and Down for line
movement and returns consumed. Anything under the composer that wants an arrow
— an @-mention list, a completion popup — therefore never sees one, because the
field has focus for exactly as long as the list is up.
Moving the hook between views on the root cannot fix this, and it was tried
twice. The chrome exposes ForgeChrome::set_prompt_key_filter, consulted inside
the field before its own handling; that is the only place that wins. Keep the
root hook as well — it carries the keys when focus is anywhere else.
The tell, if this recurs: log the events reaching the root hook. A key the field
has eaten arrives only as is_down=false (the key-up), while letters and
Tab arrive with is_down=true. And the give-away symptom is that clicking a row
makes the arrows start working, because that moves focus off the field — which
reads as the arrows being intermittent rather than as the field eating them.
set_visible(true) + request_repaint() draws the view at the size it last
had; for a view that has never been visible that is no size at all. It needs
invalidate_layout() on the view, its panel, and the ancestors.
The headless picture tests cannot catch this: render_to_png lays the tree out
from scratch on every call, so a notice that the running app could not draw
appears perfectly in the PNG. Anything about appearing has to be proven
against the running window.
macOS searches /Applications and ~/Applications, and a copy in the home
one shadows the installed copy for Spotlight and the Dock. Both this machine and
m5 accumulated an older, unsigned build there. There is no way to tell from the
running window which one answered, and a fix tested against the wrong binary
reports as not working — it cost a session on each machine. setup_m5.sh now
removes both locations; drive_app.py identifies the app by the executable
path (via ps -o comm=, since macOS pgrep has no -a) and refuses to drive
when a copy from another build is running.
Beside those sat 513 MB and 628 MB of hand-made .prev / .prev-1301 /
.backup-<timestamp> / .signed-backup bundle copies. Nothing in the repo
creates them — they are interactive "let me keep the old one" copies, and none
was ever the only record of anything, because the build directory is the
previous build and git is the history. Do not make them.
tools/rack/clean_installs.sh sweeps both kinds (dry-run by default, --yes to
remove) and setup_m5.sh sends the same script over rather than reimplementing
the rule, because two copies of a removal rule are two chances to disagree
about what is safe to delete.
An install can be complete, signed, notarized, stapled and known to
LaunchServices while Cmd-Space finds nothing — which is indistinguishable from
"it never installed", and is what one "I don't see it on m5" turned out to be.
touch the bundle, then mdimport — and the touch is the part that
matters. rsync -a preserves the SOURCE directory's mtime, and a rebuild only
changes the binary inside the bundle, so the installed .app arrives looking
older than the last index pass and Spotlight skips it as unchanged. mdimport
on its own does nothing; the tell that you are in this state is mdls reporting
null for every attribute while mdfind finds 2,300 other app bundles fine.
Indexing is asynchronous: checking immediately reported CANNOT FIND IT for
an app indexed seconds later, so the check in setup_m5.sh retries for 20s.
Query it as a structured search — kMDItemContentType == "com.apple.application-bundle" && kMDItemDisplayName == "Forge Modular" —
because mdfind -name "Forge Modular.app" finds it even when it is indexed.
tools/rack/prove_arrows.py types @br, presses Down, presses Return, and
compares captured regions. Two things make it a proof rather than a picture:
uidriver key <name|code> sends a key that produces no text — type goes
through setUnicodeString, which carries no key code, so before this there was
no way to send an arrow at all. Both need Accessibility and Screen Recording, so
they run from a Terminal on the machine, never over SSH.
patch.py / generate.py are launched with nohup + setsid so a build
survives the window closing. Every run used to redirect into one
last-run.log, and nothing stopped a second Build — so two generations
overlapping wrote over each other from offset zero. That file is not just a
transcript: the shell reads the outcome, the stage and the artifact
path out of it, so a collision puts one patch's explanation on screen beside
another patch's filename and hands Rack the wrong file. Seen on m5, two patches
finishing six seconds apart, with the log showing an acid explanation stopping
mid-word and an ambient one continuing.
Both halves are needed. The lock (start_build_with) is a build this shell
started, on a log it is watching, that has not reported an end — not
busy(), which is watching_ && outcome == running and therefore true for a
shell merely watching a log nothing has written yet; using it refused the FIRST
build of every session. The lock cannot see a run left over from a previous
launch, so the engine is also asked whether a generator process exists, and each
run claims its own log under runs/ with O_CREAT|O_EXCL — an exists() check
is not enough, because the log does not exist until the run writes it, in
another process, so two submits in the same second pick the same name.
stapler validate never says "validated"It prints The validate action worked!. A checker grepping for "validated"
reported every bundle as NOT stapled in the same run that had just stapled them.
Use the exit code. tools/rack/sign_bundles.sh signs (inner dylibs first),
notarizes, staples and re-reports all four bundles; --check reports without
changing anything, and names an ad-hoc signature as ad-hoc rather than printing
an empty authority that skims past as fine.
Forge Modular packaging must run Pulp's unattended signing doctor before its
first codesign and treat any nonzero result as terminal. Do not restore the
old PULP_SKIP_SIGNING_PREFLIGHT escape hatch or warn-and-continue: either can
fall through to the login-keychain copy and open an unanswerable GUI password
dialog. The prompt-safe path is the dedicated keychain, full partition list,
identity hash, and real timestamped probe maintained by
tools/scripts/ensure_signing_ready.sh.
For a current Forge Modular release, use
examples/forge-modular/release-package.sh against a clean detached Forge
checkout and the Release build configured from it. The script verifies all
four bundles share one exact Forge snapshot and one exact bundled Pulp
toolchain, requires tools/rack/rack_patch_decode, checks
AU/VST3/CLAP/Standalone and bundled-helper architecture, invokes the shared
signed/notarized installer primitive, and verifies the expanded PKG contains no
sibling Forge products. forge-seam/patches/BASE remains historical seam
reconstruction input and is never product provenance for this release path.
Four things could say how wide a third-party module is and, for a plugin
fetched five minutes ago, none of them does: a .vcv records no width, a
plugin.json carries none, the port map has an entry only for modules CARTOG
has scanned, and the generated .why.json sidecar copies its hp from the
inventory, which copies it from the port map. So a fresh maker's modules all
arrive at the fallback of 8 HP, and the preview drew a 30 HP sequencer squeezed
into a quarter of its width. That does not read as "we do not know how wide this
is"; it reads as panels stretched to the ceiling, which is how it was reported.
The panel SVG is a picture of the module at its true size, and every 3U panel is
128.5 mm tall, so the root tag's own aspect IS the width — in whatever units the
vendor drew it. panel_hp_from_artwork() reads it; RackModule::width_measured
is what distinguishes a measurement from the fallback, so the preview knows when
to go and look.
Two things follow:
layout_rack walks each row left to right and moves anything that would
start inside its neighbour to just after it.rack_layout
was right the whole time and the render was still wrong, because the numbers
fed into it were a guess. A geometry test that reads layout_rack's output
cannot see that. Render with the Skia backend, find the panel's ink, and check
width / height against hp * 5.08 / 128.5.modular_setting() finds the next QUOTED string after a key's colon. A JSON
number is not quoted, so asking it for one returns the FOLLOWING key's value —
confidently, in the right shape, with nothing to suggest anything went wrong.
modular_setting_int() exists for numbers. Anything numeric added to
SETTINGS_DEFAULTS needs it, and patch.py setting KEY VALUE has to convert
the argument, or the writer refuses its own choices for not being in the list.
--output-format=stream-json is refused without --verbose. Events arrive one
JSON object per line: stream_event carries the partial deltas (only with
--include-partial-messages), assistant carries whole blocks, and result
carries the finished answer — prefer it, so there is one authority for what was
said. A blocking read on a pipe cannot time itself out and a wedged call emits
no lines at all, so the deadline has to be a timer that kills the process.
ask_model() still accepts plain text on stdout: most checks here stub the CLI
with a script that prints an answer, and making each one imitate a stream to
test the code AROUND the model would be work for nothing.
Codex failures are not confined to item.completed error items. Current Codex
emits a top-level {"type":"error","message":...} followed by a
turn.failed event whose nested error.message repeats the same diagnosis.
Capture all three forms, deduplicate identical messages, and only promote them
to stderr when the process failed or produced no answer. Otherwise a real quota
or authentication refusal becomes the useless exited non-zero and said nothing, while naively appending both current events prints the same failure
twice. Pin both the current paired-event shape and the older item form in the
stream regression.
tools/rack/test_param_units.py sweeps each control's position, renders the
display value, and inverts it back. pu.from_display() returns None when the
map is not invertible — and a map with a zero-width display range is one of
those: every position renders the same string, so no position can be recovered
from it. That is a property of the control's declared display map, not a failure
of the round trip, and counting it as an offender buries the real
non-invertibility bugs under rows that were never invertible to begin with.
Classify it instead. has_no_inverse(param) answers the question from the
declared map alone, and the sweep tallies those separately from offenders. Keep
a positive check on the classifier itself — a synthetic flat map whose
from_display must return None — or the branch silently swallows every real
offender the day the predicate goes wrong.
file_lock, so it holds on Windowsgenerate.py serialises generations against one module pack with a lock file
in the temp directory. It takes that lock through tools/rack/file_lock.py
(flock on POSIX, msvcrt.locking on Windows) via claim_generation_lock();
an unconditional import fcntl made every rack tool that imports generate
fail to load on Windows. GenerationLockSafety in test_generate_safety.py
pins that a second claim exits with "already running" and that closing the
holder frees it.
まだレビューはありません。使ってみた感想をお寄せください。
概要と使いどころ
Optional AAX support for Pulp, including developer-supplied Avid SDK setup, CMake enablement, DigiShell/AAX Validator workflows, and local AAX builds on macOS or Windows.
日本語の概要は準備中です。原文の説明を表示しています。
Configure, implement, and test Pulp's optional desktop Ableton Link tempo-sync adapter while preserving the developer-supplied SDK, licensing, realtime, latency-compensation, and no-install boundaries.
日本語の概要は準備中です。原文の説明を表示しています。
Maintain Pulp's installed design-time agent capability manifest and public-surface ledger. Use when adding, removing, renaming, or materially changing public audio, MIDI, signal, timebase, or sequence APIs; registering a new algorithm for generators; changing capability support or deprecation state; or repairing agent-capabilities freshness, schema, fingerprint, tombstone, or installed-SDK tests.
日本語の概要は準備中です。原文の説明を表示しています。
Android platform development for Pulp — NDK cross-compilation, Oboe audio, Dawn/Skia GPU rendering, JNI bridge, touch interaction, emulator workflows, and end-to-end smoke validation. Covers build, deploy, debug, and the gotchas discovered during bringup.
日本語の概要は準備中です。原文の説明を表示しています。
Optional ARA support for Pulp, including developer-supplied ARA SDK setup, CMake enablement, adapter companion APIs, validation, and ARA-aware plugin implementation guidance.
日本語の概要は準備中です。原文の説明を表示しています。
The measurement surface for ALL Pulp DSP and audio-pipeline work — read it BEFORE writing or gating DSP, not only when something already sounds wrong. Covers the C++ harness (signal generators, metrics, assertions, RenderScenario, contracts), the offline Audio Doctor (magnitude/frequency response, THD/THD+N, phase/group delay), and their Python sibling the Audio Quality Lab (tools/audio/quality-lab — null residual + alignment, LTAS log-spectral distance, spectral flux/centroid, HNR, Theil-Sen drift slope, Kaiser-sinc resampling, license-guarded corpus, regression-net ratchet). TRIGGER on AUTHORING work — "build/design an oscillator/filter/synth/effect", "add a DSP module", "what should the acceptance gate be", "how do I measure aliasing / anti-aliasing / alias floor", "null against a reference", "is this DSP correct", "choose a tolerance", "golden/regression corpus for audio", "measure drift or jitter", "A/B two renders" — AND on DEBUGGING work — "is there sound / no audio / I hear nothing", "does this filter/compressor/synth/delay produce the right signal", "prove the DSP / prove the contract", "measure the frequency response", "what's the THD / is it distorting", "what's the group delay / phase response / measured latency", "magnitude response curve", "render a test tone and assert", "audio regression", "64-frame works but 128 is silent", "sample-rate change pitch-shifted it", "describe what's in this buffer", "audio doctor", "compare before/after a DSP refactor". Reach for this BEFORE hand-rolling any FFT, null test, alias measurement, pitch tracker, or golden-render script — most of it already exists in one of the two lanes. Test/tool layer over HeadlessHost — deterministic, no audio device, no speakers. Off the realtime thread entirely.
日本語の概要は準備中です。原文の説明を表示しています。