aax
無料Optional AAX support for Pulp, including developer-supplied Avid SDK setup, CMake enablement, DigiShell/AAX Validator workflows, and local AAX builds on macOS or Windows.
日本語の概要は準備中です。原文の説明を表示しています。
Maintain Pulp's installed design-time agent capability manifest and public-surface ledger. Use when adding, removing, renaming, or materially changing public audio, MIDI, signal, timebase, or sequence APIs; registering a new algorithm for generators; changing capability support or deprecation state; or repairing agent-capabilities freshness, schema, fingerprint, tombstone, or installed-SDK tests.
インストール方法を見るインストールする前に、エージェントに与えられる指示の中身を確認できます。
Maintain three related artifacts:
agent-capabilities.json is the installed consumer contract: curated keys,
versions, digests, status, evolution, typed C++ bindings, and partial-coverage
semantics.agent-capability-surface.json is the maintenance ledger: every public header
in the covered roots, its byte fingerprint, and its reviewed disposition.tools/agent-capabilities/contract-history.json is the repository-only,
append-only evolution history checked against the protected Git tip. Shallow
GitHub Actions checkouts fetch the immutable event base SHA when necessary.The consumer manifest, its schema, and the handoff schema install into the SDK.
Official release packaging stamps agent-capability-handoff.json only after
installation; it binds the exact SDK source SHA and platform to the installed
importer's SHA-256 plus the installed manifest's exact content and byte hash.
The release archive verifier must require and revalidate that identity at the
configured capability-handoff floor. The surface ledger, surface schema, legacy
baseline, and contract history are maintenance artifacts and must not be
installed.
When adding an installed SDK library in PulpInstallRules.cmake, register its
archive stem in release_product_matrix.json and classify every newly covered
public header here in the same change. A successful CMake export alone does not
prove the release archive or installed agent-capability contract is complete.
Keep both separate from the unified runtime control platform. This contract may describe what an SDK can design or generate; it must never contain runtime operations, grants, policy, risk decisions, instances, activation, sessions, revocation, or receipts.
pulp authority list|query is the routing-only navigator across this and the
other bounded machine authorities. For design-time capability questions it
must point to agent-capabilities.json and this manifest checker; it must not
copy capability rows into authority-navigation.json, treat an absent key as
unsupported, or infer any live state. In a source checkout it binds its answer
to the checkout registry. In a downstream project it resolves the exact SDK
selected by pulp.toml without materializing it; otherwise an SDK-installed
CLI may bind to its adjacent share/pulp/authority-navigation.json.
Source-only routes report unavailable, and a missing selected SDK never falls
back to another installed version.
The standalone CLI-only release archive has no SDK registry and must fail
clearly rather than search another checkout or SDK. A reported
query_or_validator is descriptive guidance only and is never executed by the
navigator.
The installed SDK also ships canonical runtime-control headers, CMake helpers,
and control-authoring examples. Their presence in the same install tree does
not make them agent-capability rows: keep runtime-control operations and policy
out of agent-capabilities.json, and keep the capability surface ledger focused
on the design-time public-header contract.
The A3 GPU startup-health surface is a concrete example. The installed
ControlGpuHealthProvider, ControlGpuHealthViewAdapter,
ControlGpuHealthReadExecutor, and pulp.gpu-health-read-result.v1 types are
runtime control plumbing. Do not add dev.pulp.gpu/health.read@1, its grants,
instances, receipts, measurement campaigns, or B4 disposition to
agent-capabilities.json. A public-header ledger classification, where one is
required by the covered roots, records only that the header was reviewed; it
must not turn this runtime operation into a generator-facing design capability.
The checkout-only gpu_first_visible_a3_campaign.py runner is also runtime
acceptance tooling, not an installed SDK capability: do not add its adapter
request/receipt schemas, 10+10 lifecycle ledger, or source-binding receipt to
the design-time capability catalog.
The same classification holds when that surface grows a producer.
FrameObservation::gpu_submission_observed is now fed by two WindowHost
queries (supports_gpu_submission_evidence(),
last_frame_gpu_submission_observed()). Those are public installed-header
additions and get a public-header ledger classification like any other, but
they stay OUT of agent-capabilities.json: a host telling the inspector whether
its last frame reached an output is runtime control plumbing, not a
generator-facing design capability, and a plausible WindowHost:: prefix is not
evidence otherwise.
A stopped-install runtime observer can change a reviewed header without adding
an advertised design capability. Refresh the affected fingerprints and inspect
the generated manifest diff: its legacy method projection can change even when
all curated capability contracts stay identical. Bind related pending exposure
evidence to the specific API or header fingerprint, not a literal global
SURFACE_INVENTORY_VERSION; an unrelated inventory update must not invalidate
that evidence. None of these source checks establishes installed execution.
For a new public header or symbol:
EXPORTS in the domain-appropriate
tools/scripts/agent_capability_catalog_*.py module, add typed bindings for
every advertised entrypoint/operation, and record the current header
fingerprint. Add a nonempty _link_probe that constructs or invokes the
real API rather than merely taking sizeof, and start a new key at contract
version 1.0. agent_capability_manifest.py assembles those catalogs; do
not put capability rows back into that orchestrator.tools/scripts/agent_capability_registry.py instead:
capability_support, infrastructure, or unsupported_capability. Give a
durable rationale. Never grow the frozen legacy_unreviewed baseline.SURFACE_INVENTORY_VERSION for any ledger change. Increase
MANIFEST_REVISION whenever the installed manifest changes.For overloaded C++ free functions, keep the public qualified_name as the real
API name and give each overload a distinct binding role. Use private generator
metadata for an explicit static_cast address expression so the generated
compile fixture proves the intended signature without leaking fixture syntax
into the installed contract. Each overload still needs its own operational
probe with arguments that select and invoke that overload.
If generation reports that a new public header is unclassified or has no
covered public target owner, do not retry --write or add a blanket exception.
Add the curated capability or reviewed disposition first, register the exact
minimal owner in REVIEWED_MINIMAL_TARGETS, and then regenerate. A capability
binding alone cannot establish which installed CMake target owns its header.
For a non-static member-function binding, keep the public qualified_name as
the real class-qualified method, provide an exact pointer-to-member
address_expression, and use an explicit-object member_function_call probe.
The generator must retain address references with auto volatile; auto *
cannot represent a pointer-to-member. The installed-SDK suite runs the matching
operational probe independently for every binding, not only in the aggregate
capability consumer.
A new TSP algorithm is therefore detected automatically but not advertised by guesswork: the new/changed public header fails the ledger gate until its owner makes the explicit registration or non-capability classification.
For a fixed-capacity record algebra such as music.pattern-development, bind
the stable record, error, configuration, and result types as well as every
advertised free function. Each free function needs its own operational probe;
a type-only row or one aggregate probe cannot establish that installed
consumers can execute density, fill, set-algebra, ID, and morph operations.
Keep scheduling, clocks, note ownership, and publication outside this manifest.
For bounded sampled-target FIR design, register the public
pulp/signal/fir_design.hpp entry point as signal.fir-design, keep it
offline-only, and have the generated compile fixture invoke an empty-target
request in addition to taking the exact function pointer. This preserves the
contract's proof that the published binding is operational rather than merely
type-visible.
For a bounded event-domain MIDI player such as midi.linear-step-player
(pulp/midi/step_player.hpp), follow the midi.arpeggiator row shape: one
cpp_type binding on the template entrypoint with a member_call probe that
default-constructs it and invokes a const query, and add the new key to the
capability_keys list of each support header the kernel actually includes
(e.g. utility_contract.hpp), not to every MIDI support header.
For an existing capability change:
Update the reviewed header fingerprint for every public-header byte change,
even when the consumer contract is unchanged. Increase the surface inventory
version. --write cannot do the first of those for you: it reports the
measured digest and exits nonzero, because a generator free to restamp a
fingerprint would silently launder every unreviewed header edit. Paste the
measured digest over the declared one in agent_capability_registry.py, bump
SURFACE_INVENTORY_VERSION in agent_capability_manifest.py, then --write.
Exception: when the only stale fingerprint is a catalog binding's
header_fingerprint (e.g. agent_capability_catalog_performance.py, which
pins processor.hpp for process_block) and no capability's surface moved,
paste the measured digest there and --write WITHOUT bumping
SURFACE_INVENTORY_VERSION: the bump is refused with "surface
inventory_version changed without a surface change". A Processor virtual
that no capability exposes (an editor hook, say) is exactly this case.
Before choosing the next SURFACE_INVENTORY_VERSION, check whether an open
branch already claimed it. Two branches that both bump 79 to 80 do not
conflict — the edits are identical, so git merges them silently and the
surface ends up recorded under a version that describes two different
inventories. Nothing downstream catches that, because each side passes
--check on its own.
for p in $(ghapp api "repos/Generous-Corp/pulp/pulls?state=open&per_page=100" --jq '.[].number'); do
ghapp api "repos/Generous-Corp/pulp/pulls/$p/files?per_page=100" \
--jq '.[] | select(.filename=="tools/scripts/agent_capability_manifest.py") | .patch' 2>/dev/null \
| grep -q "^+SURFACE_INVENTORY_VERSION" && echo "PR #$p claims a version"
done
Skip past any claimed integer rather than racing for it; the values only have
to be distinct and increasing, not contiguous. ghapp pr list --state open --json number,files answers "does any open PR touch the manifest file" in
one call; filter it before fetching any patch. A stacked branch (one PR
based on another) takes the next integer again: the base PR claims N+1, the
stacked one N+2, and each repoints any pending needle that pins the integer.
A new header under a covered root also moves the consumption census
(docs/status/consumption-profiles.json; test_consumption_census.py
HeaderNameDrift). That test reads names from the git index, so it keeps
failing until the header is tracked. Regenerate with
consumption_census.py --build-dir <this checkout's build> --write; pointing
--build-dir at a sibling worktree's build rewrites every include root to
the other tree, so for a stacked branch without its own build add the one
header line by hand instead.
A pending sequencer-exposure row can pin the current inventory integer as an
evidence needle ("SURFACE_INVENTORY_VERSION = 95"). Bumping the version
stales that row and sequencer_exposure_check.py fails on it, not on your
change. Repoint the needle to the new integer — it is a pending row, so the
edit is allowed — and grep every pending row for the outgoing literal first.
A summary-only edit to a capability row is not a contract change: raising
contract_version for it is refused ("contract_version changed without a
contract change"). Leave the version, bump MANIFEST_REVISION, and let
--write regenerate the installed manifest.
A byte change to a header still classified legacy_unreviewed cannot be
repaired by restamping tools/agent-capabilities/legacy-unreviewed-baseline.json.
FROZEN_LEGACY_COUNT and FROZEN_LEGACY_DIGEST in
agent_capability_surface.py pin that file, so editing a fingerprint there
and moving the constants to match is exactly the laundering the frozen
baseline exists to prevent. Graduate the header instead: add a reviewed
disposition for it to agent_capability_registry.py with the measured
fingerprint and a durable rationale, and leave the baseline untouched. The
reviewed branch is consulted ahead of the baseline, so the stale baseline
entry is simply no longer read; dozens of headers already sit in both.
--write also appends a full snapshot to
tools/agent-capabilities/contract-history.json — tens of thousands of lines
that dwarf the change that caused them. --check does not require it, so for
a byte-level fingerprint refresh, revert that file and keep the three-file
change; --check still reports fresh. Reverting it is not free of meaning,
so keep the snapshot when the change is a real contract movement whose history
someone will read back.
Increase the capability minor version for compatible additive contract changes.
Increase the capability major version when a binding is removed, renamed, or replaced, or when lifecycle, RT, state, seed, domain, units, latency, tail, or scheduling semantics change incompatibly.
Keep seed_model and determinism separate. When
determinism-contract-v1 is required, every live row must declare
repeatability, block-partition behavior, platform scope, and whether transport
history is an input. Strengthening a determinism promise is additive;
weakening or removing one requires a major increase or a new successor key.
Treat a minor-0 row with no determinism as unspecified. Consumers that
require determinism must reject it and must reject unknown required features.
Leave the capability version unchanged for summary-only wording. The generated digest excludes the summary but covers the material contract.
Keep numeric parameter ranges/defaults/choices in forge-catalog.json; use a
forge_descriptor reference instead of copying them.
For realtime capability implementations, do not treat a short-range
std::stable_sort as allocation-free merely because the macOS/libc++ probe is
green. libstdc++ may allocate scratch space for every non-empty range while
libc++ keeps small trivially-copyable ranges in place. Prefer a bounded
in-place stable ordering algorithm when the capability already declares a
fixed maximum, and size the allocation negative control beyond libc++'s short
in-place threshold so either standard library can expose a regression.
For removal:
status: deprecated with a matching
deprecated evolution state and ordered lifecycle versions. A capability that
is active in the protected base may not be removed in the current change.status: removed capability tombstone
with its introduction/deprecation versions, last version, and digest before
deleting the row. Never reuse a tombstoned key.The published capability table now includes bounded, versioned registered clip content and trusted note-renderer hooks. Keep that row aligned with the compile contract: note output/reset state only, a 4096-note fragment cap, and explicit refusal of trimmed nesting and nondefault-production wire serialization.
status is the explicit support claim. Use only
stable, usable, experimental, partial, unsupported, or deprecated;
never publish planned work.unsupported_capability is an explicit negative claim.legacy_unreviewed means only that no machine-readable claim has been
reviewed yet.partial, so an absent key means unknown, not
unsupported.signal.spectral-mask-processor is the shared streaming STFT/WOLA layer for
products that apply authored spectral gain tables. Installed-SDK consumers
include <pulp/signal/spectral_mask_processor.hpp> and link Pulp::signal.
Prepare it off the audio thread, publish layouts or compiled tables from a
control thread, and call process() or process_frame() on the audio thread.
The processor owns frame-boundary table adoption, gain interpolation,
overlap-add reconstruction, latency reporting, and latency-aligned dry/wet.
Sample-scheduled host automation is the bounded exception to control-thread
publication. A single audio owner may call set_layout_rt() with the latest
fixed-capacity layout while applying events at their block offsets. The
processor copies that layout into prepared storage and compiles/adopts it only
at the next spectral-frame boundary, without allocation, locks, or a
control-thread round trip. Do not call set_layout_rt() from multiple writers
or use it as a replacement for UI/state-restore publish_layout(); the two
paths deliberately keep separate writer contracts and converge only at the
audio owner's frame-boundary adoption point.
Use categorical mask entries for true mute: a muted bin is multiplied by exact zero, not represented by a finite decibel floor. Reuse this processor for zoomable filter banks, spectral gates, freezes, morphing, and related products instead of rebuilding an application-local STFT lifecycle or publication protocol. Analyzer snapshots and captured-frame storage are separate layers; do not infer them from this capability or duplicate them inside the processor.
Before presenting every authored band as independently controllable, call
analyze_spectral_band_resolution() with the product's layout, sample rate,
and FFT size. Its fixed-capacity report counts directly owned viewport bins per
band and excludes exterior edge-band extension. fully_represented() == false
means the UI or profile selector must disclose the resolution limit, select a
higher supported geometry, or use a different filter architecture; zoom alone
cannot create additional FFT bins.
pulp::view::VisualizationBridge is the shared realtime-safe audio-to-UI tap
for spectrum, waveform, and meter consumers. Configure it while fully
quiescent, call process() from the audio callback, and give exactly one UI
thread ownership of poll() plus the snapshot reads. The callback path only
meters and copies into fixed SPSC storage; FFT and waveform assembly happen in
the bounded, non-realtime poll() call.
read_spectrum() and read_waveform() remain cheap snapshot reads for source
compatibility. They do not analyze newly captured audio. A consumer that needs
fresh data must schedule poll() first, then read or use the explicit
peek_spectrum() / peek_waveform() aliases. Treat capture overflow, rejected
channel topology, and positive-length missing-channel callbacks as continuity
breaks: the bridge advances its epoch and never joins audio across the gap.
Keep configure() and reset() quiescent; neither is concurrent with the
audio producer or UI consumer.
Do not use the bootstrap or unpublished-migration switches during normal work.
When the base branch has already advanced either counter, recompute the next
MANIFEST_REVISION and SURFACE_INVENTORY_VERSION from that exact base before
running --write; replaying stale projection counters can silently reuse an
already-published contract identity.
Use Python 3.10 or newer: transactional generation uses modern standard-library
APIs such as zip(..., strict=True), so an older system python3 can fail before
validating the contract. Run installed-SDK capability tests from a Release build.
A Debug/coverage build is an invalid positive control because the installed SDK
guard intentionally refuses unacknowledged Debug SDK consumers. Also keep the
build tree free of stale nested SDK install prefixes: archive-mutation checks
require one owning build-tree library per target, and an old consumer-smoke
prefix/lib can create a false duplicate-owner failure. Move such generated
fixtures aside and rerun the same test before changing capability code or
weakening the archive check.
Run:
python3 tools/scripts/agent_capability_manifest.py --write
python3 tools/scripts/agent_capability_manifest.py --check
python3 tools/scripts/test_agent_capability_manifest.py
cmake --build build --config Release --target pulp-test-agent-capability-compile
ctest --test-dir build -C Release -R '^agent-capability-' --output-on-failure
The installed-SDK test must install to an isolated prefix, verify maintenance
artifacts are absent, read the installed schema and manifest, and independently
compile/link/run every capability and every typed binding against only its
declared minimal target. The positive consumers may share one CMake configure
and bounded parallel build, but each proof must retain its own source, executable
target, declared-minimal-target link, CMake File API isolation inspection, and
process execution. Coverage validation must reject missing, duplicate, or wrong-
target proofs before configuration. The test must reject wrong-target
declarations and checkout-path leakage, and use configuration-aware build/install
and executable paths. Because that proof can take roughly 18 minutes when every
consumer is configured and built serially on an Apple runner, its CTest
registration carries slow-affected;agent-capability-installed-sdk (every
lane's slow exclusion matches it). Ordinary PR and merge-group corpora exclude
it, but classify_changes.py restores the exact
test on the parallel macOS and Linux matrix legs when the diff touches capability manifests,
schemas, history, registries/generators, vocabulary, install rules, or their
compile tests. All CMake target/export definitions are included because an
INSTALL_INTERFACE, exported dependency, or target-name change can break the
isolated consumer even outside PulpInstallRules.cmake. A selected documentation
surface also forces allocation of the containing native job; unknown or
unavailable diffs run it fail-closed. Keep that
affected-surface step (macOS is the required queue context; Linux preserves its
platform-specific export proof) and the unfiltered main/nightly proof when changing its
label or registration; slow alone is not authorization to stop enforcing it.
When checking CMake File API include paths, permit paths outside the install
prefix only when CMake marks them as system includes; transitive platform and
third-party headers may be legitimate, but non-system source/build leakage is
still a failure. When finding build-tree archives for mutation controls, exclude
the staged install prefix because retrying the test leaves installed archives
under the build directory.
The official-SDK handoff self-test separately covers exact identity plus wrong
source SHA, importer hash, capability hash, and schema-invalid documents.
test/cmake/quality_tests.cmake is also the registry for capability-manifest
and adjacent policy self-tests. When adding a new Python policy test there,
register the test explicitly in the same change; merely creating a
tools/ci/test_*.py file does not make CTest execute it.
That registry also runs the browser DPR adapter self-test. Keep its exact Playwright-version, artifact-confinement, product-digest, typed-metric, timer calibration, logical-input, and same-content fidelity negatives registered; an executable measurement script without the CTest entry is not maintained evidence tooling.
That registry also carries the trusted Vellum merge self-test. Keep its clean base+head positive control and real content-conflict negative control together: the required gate must prove it can construct the exact two-parent candidate, and that the same constructor refuses a conflicted candidate before validation.
The surface fingerprint is intentionally conservative SHA-256 over full header bytes. Do not weaken it with regex symbol extraction. A future pinned-Clang AST inventory may reduce comment/private-detail churn only if its version and toolchain are pinned and mutation tests retain add/remove/change detection.
Because the fingerprint covers full bytes, editing only comments in a
capability header is a surface change and --write will refuse it twice
before it succeeds. The declared fingerprint lives in the catalog source, not
just the generated JSON, so the order is: edit the header, then replace every
occurrence of the old digest in the owning agent_capability_catalog_*.py
(one per binding, so a single header can hold a dozen copies), then raise
MANIFEST_REVISION and SURFACE_INVENTORY_VERSION, then run --write once.
Derive both counters from the CURRENT protected base every time. A capability transaction can land while yours waits in the merge queue, which takes the numbers you reserved and leaves your branch conflicting on exactly those two constant lines. Re-read them from main and regenerate rather than resolving that conflict by hand.
Do that with the tool, not by hand:
python3 tools/scripts/agent_capability_rederive.py --print # read-only: what would change
python3 tools/scripts/agent_capability_rederive.py # rewrite both, then --write
It resolves the same protected tip --check uses, compares this tree's
generated material against the base's with the counter excluded, and moves each
counter only when its own material actually moved. That asymmetry is the part
worth not hand-rolling: over-bumping is not the safe direction. Advancing a
counter whose material is identical fails the opposite rule,
... changed without a manifest change, so "bump both to be safe" trades one
red gate for another.
A taken counter does not always announce itself as a conflict. Re-read both
counters after every merge of the protected base, not only after git reports
one. When two lanes reserve the same next integer, the two sides hold
character-identical constant lines, so the merge is clean and silent — and the
increase is what gets annihilated: the surface has changed (your fingerprint)
while the counter equals the base again. That surfaces much later as
STALE: public surface changed without an inventory_version increase, on every
platform at once, naming generated files the diff appears not to touch.
Recovery needs the surface document reset first. --write derives from the
on-disk artifact, which already carries your fingerprint at the taken counter,
so raising the counter alone fails the opposite rule instead:
git checkout origin/main -- docs/status/agent-capability-surface.json
python3 tools/scripts/agent_capability_rederive.py
Reset that one document and nothing else. The same digest also lives in
REVIEWED_HEADERS in tools/scripts/agent_capability_registry.py, and that copy
must keep the NEW value — resetting it too restores the stale digest and
reproduces the original failure. No integer is picked by hand: rederive.py
resolves the protected tip, so it lands on whatever is free.
It refuses rather than guesses when the surface has unresolved problems — a changed header with a stale fingerprint has no stable material to derive from, and its fingerprints must be refreshed first. Counters are decided LAST.
Re-running it is idempotent: the counter a tree currently holds is not evidence of anything, so a tree already carrying a stale reservation is derived back down to what its material justifies.
A reviewed public header's fingerprint can be declared in either
agent_capability_registry.py (as a REVIEWED_HEADERS entry) or a
agent_capability_catalog_*.py module (as a binding's header_fingerprint).
--check names the header and both digests but not the file, and grepping the
registry for a header that is declared in a catalog finds nothing — which reads
as "not tracked" rather than "declared elsewhere". Find the declaration by the
expected digest, which is unique:
grep -rn "<the expected sha256 from --check>" tools/
tools/agent-capabilities/contract-history.json will also match; it is the
appended snapshot, never the declaration. Measured on
pulp/timebase/grid_projection.hpp, whose digest lives in
agent_capability_catalog_timing.py as the project_grid binding.
--check for you — but only against a resolved basegates.sh and .githooks/pre-push both run --check before a push, whenever
the diff touches an installed public header or a capability registry/manifest
path. Both pass PULP_AGENT_CAPABILITY_BASE_REF set to a resolved commit,
never the base's name, because a bare run resolves the merge base while CI
forces the base tip, and those disagree: measured on one tree in one second,
80 against the merge base and 82 against the tip — and 80 was exactly the
number a branch had already carried into the merge queue. Reproduce a CI
verdict the same way rather than trusting a bare local run.
Two things the gate deliberately does not hide:
SKIPPED and says a skip is not a pass.
Nothing green is implied by silence there.This matters more than the other push gates because the failure is
unrepairable after the fact. The merge group is the only place that sees a
stale counter, and a pull request that has entered the merge queue refuses a
push with GH006: … Branches that are queued for merging cannot be updated —
so the fix is locked out by the same queue that is about to reject the branch.
Removing it from the queue is itself gated (queue-removal-guard: refusing unaudited merge-queue removal), which leaves waiting for eviction as the only
unprivileged route. Fifteen seconds before the push replaces roughly an hour of
merge-group time plus a queue lockout.
Commit the merge before you run it. During an uncommitted merge the incoming
tip is not yet an ancestor of HEAD, so base resolution steps back to the merge
base — a commit that predates both sides — and the counter derived from it is
stale while looking entirely plausible. Observed live: it read base 28/45 while
main already held 30/47. The tool now refuses in that state rather than
answering the wrong question, so the sequence when resolving a conflict is:
resolve → commit the merge → refresh fingerprints → re-derive → --write.
Take the base's side of a contract-history.json conflict. Merging a base
that landed its own transaction conflicts here as a both-append: your side holds
the entry your --write appended, the base holds every entry up to its own tip,
and the two are not reconcilable line by line. git checkout --theirs is right,
and it does not discard your transaction — updated_history_entries() appends
the entry describing the previous manifest, so history never carries your
branch's own state to begin with. That state lives in
docs/status/agent-capabilities.json, which the re-derive regenerates
afterwards. Hand-merging the two runs instead produces a history whose tail is
not the protected base's entry, which is exactly what the append-only check
rejects.
Regenerate exactly once from final header bytes. Each --write appends a full
entry to contract-history.json, so editing the header again after a
successful --write leaves two entries for one logical change.
--write is not idempotent, so never use it to verify a transaction. A
second --write on an unchanged tree still appends a second history entry, and
it reports the same cheerful wrote ... line either way — so the check reads as
confirmation while it is the thing creating the defect. The reviewer then sees
capability history is not append-only relative to the protected base for a
transaction that was correct until it was checked. Verify with --check, which
writes nothing and answers the same question (fresh; N keys and M public headers checked). If a verification --write already ran, do not try to prune
the entry by hand: reset the four generated artifacts to the protected base and
regenerate once.
Prove the append is single rather than assuming it. git diff --numstat showing
zero deletions is necessary but not sufficient — two appends are also
deletion-free. Compare entry counts:
python3 -c "
import json,subprocess
c=json.load(open('tools/agent-capabilities/contract-history.json'))['entries']
b=json.loads(subprocess.run(['git','show','origin/main:tools/agent-capabilities/contract-history.json'],
capture_output=True,text=True,check=True).stdout)['entries']
print('delta', len(c)-len(b), 'prefix-identical', c[:len(b)]==b)"
--check's diagnostics go to stderr, not stdout. A script that scrapes the
expected/got fingerprint pairs from captured stdout alone silently sees nothing
and reports success.
--write validates one precondition at a time and stops at the first failure,
so a new-header capability surfaces as four unrelated-looking errors in
sequence rather than one checklist. Make all four edits before running it:
capability(...) block in
agent_capability_catalog_<domain>.py."pulp/<domain>/<new>.hpp": "Pulp::<domain>",
in agent_capability_registry.py. Without it the error is
bindings[N] include has no covered public target owner, which names the
binding rather than the missing map entry — the message points away from
the fix.<domain>.hpp changes that
umbrella's digest too. Update the declared value to the got sha256: the
error reports. The stale digest also appears in contract-history.json;
do not edit those — history is append-only.MANIFEST_REVISION and SURFACE_INVENTORY_VERSION,
reported as two separate errors.--writeprevious is the on-disk docs/status/agent-capabilities.json, not the base
ref. So a tree where a previous --write partially succeeded, or where
rederive has reset some artifacts and not others, makes the gate compare
against a state that never existed. The symptom is a pair of contradictory
verdicts across consecutive runs on an unchanged diff:
agent-capabilities: INVALID: <key> changed without a contract_version increase
agent-capabilities: INVALID: new capability must start at contract version 1.0: <key>
Both cannot be true. Reading either as the real constraint sends you to invent a
version number. The fix is to make previous the base again and write once:
git checkout origin/main -- docs/status/agent-capabilities.json \
docs/status/agent-capability-surface.json \
tools/agent-capabilities/contract-history.json \
test/test_agent_capability_compile.cpp
python3 tools/scripts/agent_capability_manifest.py --write
contract-history.json is safe for a NEW capability, not for a changed oneDropping the ~18k-line history snapshot keeps a capability PR reviewable, and it
is correct when you are only introducing a capability. It is wrong on a
change that moves an existing capability's contract_version: --check reads
the history as its new-capability oracle, so without an entry the capability
reads as new on every subsequent run and must start at contract version 1.0
can never be satisfied — the version is then unbumpable.
Test it rather than assume: run --check with the snapshot retained and again
with it dropped. If the dropped run fails must start at contract version 1.0,
the snapshot is load-bearing for your change and has to ship.
One rule worth knowing before you reach for a version bump at all:
_binding_identity (agent_capability_evolution.py) excludes
header_fingerprint, and bindings is popped before the non-binding
comparison — so repointing a digest is NOT a contract change. Any edit to a
non-binding field such as output_domain IS, and the gate classifies every such
edit as breaking, so the MAJOR moves, never the minor.
agent_capability_rederive.py first restores the generated manifest, surface and
history from the protected base, then calls --write. Its output can print the
write before the reset message because the subprocess output is flushed first;
that display order does not mean the reset happened last.
updated_history_entries() in agent_capability_history.py appends the
previous manifest/surface snapshot unless it is already the last entry. It
stores the current snapshot directly only for an empty/bootstrap history. The
new contract and header fingerprints belong in the current generated manifest
and surface; their absence from the last history entry is not a failed write.
After rederive, verify the protected history is an unchanged prefix, inspect
which prior snapshot was appended, and run agent_capability_manifest.py --check.
Do not blindly call --write again to put the current fingerprint in history:
that second call treats the just-generated artifacts as previous material and
adds an unnecessary current snapshot. A one-entry delta is expected when the
protected base's own snapshot was not already recorded, but check the material
rather than prescribing a count independently of the starting history.
rederive.py refuses while a merge is in progress when the incoming commit is
not the protected base. Commit the merge before deriving counters so the base
resolver cannot silently compare against an older merge base. The checker's
protected-base prefix and evolution rules remain authoritative; never repair
history by rewriting or deleting protected entries.
core/host header is invisible until it is NAMEDcore/host/include/pulp/host is not one of PUBLIC_ROOTS. Every other
domain is discovered by rglob, so a new public header there is fingerprinted
the moment it exists; a host header reaches the surface only if it is listed
by name in REVIEWED_HOST_HEADERS (agent_capability_surface.py), whose records
carry REVIEWED_HOST_DOMAIN (host).
The consequence is easy to miss because it reads as success: adding an installed
header under core/host/include leaves --check reporting fresh, since the
surface never saw it. The header ships public and unguarded, and a later byte
change to it trips nothing. Add it to the tuple in the same slice that adds the
header.
Two further rules, both learned by getting them wrong:
REVIEWED_HEADERS row cannot rescue an undiscovered header.
discover_headers() builds the current map first, and REVIEWED_HEADERS is
validated against that map, so a row for a host header missing from the tuple
fails reviewed public header is missing rather than registering it. Name it in
the tuple first; only then does a row (or a catalog binding) resolve.If the header is bound by a capability in a catalog module, it needs no reviewed row at all — see the next section.
REVIEWED_HEADERS rowA header named by a binding(...) in a catalog is already a capability
entrypoint. Adding a capability_support row for it in REVIEWED_HEADERS fails
with headers cannot be both capability entrypoints and separately reviewed,
and that failure arrives from rederive.py as a surface problem, which reads
like a fingerprint issue rather than a duplicate-registration one.
For a new kernel header, REVIEWED_MINIMAL_TARGETS is the only registry edit it
needs. REVIEWED_HEADERS rows are for headers that no catalog binds: shared
vocabulary headers (whose capability_keys list names the kernels expressed
over them) and infrastructure headers such as a private detail/ helper,
which bind no key of their own.
header_fingerprint IS the SHA-256 of the header file's bytesThe surface document says so itself — "fingerprint_algorithm": "sha256-file-bytes" — and every one of the 453 declared fingerprints agrees
with the file on disk:
import json, hashlib, pathlib
d = json.load(open('docs/status/agent-capability-surface.json'))
differ = [r['source'] for r in d['headers']
if 'sha256:' + hashlib.sha256(pathlib.Path(r['source']).read_bytes()).hexdigest()
!= r['fingerprint']]
print(len(d['headers']), 'headers,', len(differ), 'differ') # 453 headers, 0 differ
The only real difference is the sha256: prefix the declared value carries and
bare shasum output does not — which is what makes the two look unequal at a
glance. Corrupt one declared value and the same loop reports it, so a clean run
is a measurement rather than a tautology.
An earlier revision of this section claimed the two "legitimately differ" and told you not to reconcile them by hashing the file. That was wrong, and it is the expensive kind of wrong: it reads as permission to paper over a genuine mismatch, when a mismatch means the header moved and the surface did not.
--write cannot fix a fingerprint — it is authored, not generatedThe fingerprint is declared in two places that must be edited by hand and kept
equal: the header_fingerprint= literal in the capability's EXPORTS row in
tools/scripts/agent_capability_catalog_performance.py, and the fingerprint
field for that source in docs/status/agent-capability-surface.json.
Regeneration checks them; it does not author them. So after changing a public
capability header:
shasum -a 256 <header> and prefix the digest with sha256:,SURFACE_INVENTORY_VERSION in tools/scripts/agent_capability_manifest.py
(surface axis — see the next section: this costs no contract bump),--check and confirm it reports fresh.--write adds a fourth fileThe recipe above names two. A third holds the value the checker actually
compares against: the fingerprint field of the header's row in
REVIEWED_HEADERS in tools/scripts/agent_capability_registry.py. Updating
only the surface document leaves --write and --check both reporting the
same mismatch with the old digest, which reads as "regeneration is broken"
rather than "one more literal to edit". Grep the old digest across
tools/ and docs/ and replace every hit.
--write then also appends a full manifest snapshot to
tools/agent-capabilities/contract-history.json — tens of thousands of lines
recording the (manifest_revision, inventory_version) pair. That append is
not required for freshness: --check reports fresh on the three-file edit
alone, so discard the history hunk unless the change is one whose lineage the
history is meant to carry. Confirm with --check rather than assuming either
way.
A binding's identity is (role, kind, include, qualified_name, target, availability). header_fingerprint is not a component, and it does not
appear in the capability row at all — it lives only in the surface document,
versioned on its own SURFACE_INVENTORY_VERSION axis. So a pure header-bytes
change is a surface-axis event that is invisible to every capability contract
payload.
Two agents independently reasoned "fingerprint ∈ bindings ∈ contract_payload,
therefore this needs a version bump," each having verified _binding_identity
(which legitimately excludes the fingerprint) and let that stand in for
checking the snapshot's actual binding shape one level down. The generator
rejected it with contract_version changed without a contract change. If you
are adding a function to a header that already backs a capability, expect the
existing key to stay at its current version and the change to be absorbed by
the two counters.
--write / --check verify each binding's header_fingerprint against the
file's real bytes, and several capabilities may bind into one header. Adding a
second entry point to an existing header therefore breaks the FIRST capability's
fingerprint too, and the failure names the header rather than the capability, so
it is easy to scope the fix too narrowly:
capability bindings disagree on fingerprint: pulp/signal/fir_design.hpp
public header fingerprint changed: pulp/signal/fir_design.hpp; expected <old>, got <new>
Refresh EVERY binding that pins that header, not only the one you added. Compute the value from the post-edit bytes:
python3 -c "import hashlib;print('sha256:'+hashlib.sha256(open('<header>','rb').read()).hexdigest())"
Then run tools/scripts/agent_capability_rederive.py. It decides the counters
LAST and refuses to run while the surface has unresolved problems, saying so
explicitly. That ordering is deliberate: fix fingerprints first, derive counters
second. Never hand-edit manifest_revision / inventory_version in the
generated JSON; they are projected from source constants and are regenerated.
legacy_unreviewed header is a capability transactionAny edit to a public header sitting in the frozen legacy bucket fails
--write/--check with public header fingerprint changed, and you cannot fix
it by updating the fingerprint in
tools/agent-capabilities/legacy-unreviewed-baseline.json: the next error is
legacy baseline digest changed; the frozen legacy_unreviewed set may only shrink through explicit reviewed classifications. That is deliberate. The frozen set is
content-pinned so headers in it cannot be edited silently.
The sanctioned path is to classify the header OUT of the bucket. That takes two edits, and the baseline file is not one of them:
REVIEWED_HEADERS in tools/scripts/agent_capability_registry.py
with its NEW fingerprint, a disposition, and a rationale; thenpython3 tools/scripts/agent_capability_rederive.py, not a hand-edit, to
move the counters. Editing manifest_revision / inventory_version directly
in the generated JSON does nothing: they are projected from
MANIFEST_REVISION / SURFACE_INVENTORY_VERSION constants, so --write
regenerates them and still reports changed without a revision increase.Do NOT delete the entry from the baseline, decrement frozen_count, recompute
entries_digest, or touch FROZEN_LEGACY_COUNT / FROZEN_LEGACY_DIGEST. The
declaration alone satisfies the fingerprint check; the surface document derives
each header's disposition from REVIEWED_HEADERS first, so a declared header
stops being counted as legacy without the snapshot changing at all. Editing those
pinned constants to make a PR pass removes the deliberate guard — see "A header
in the frozen legacy baseline does NOT require unfreezing anything" below, which
is the authoritative statement. Afterwards confirm the baseline file has no diff
and its entry count is unchanged; --check should report fresh, and the
surface's legacy_unreviewed count should have dropped by exactly the number of
headers you declared.
Only classify a header when the classification is already defensible from a written decision, and only when the edit that tripped the gate is load-bearing for the change (the triage in that later section). Inventing a classification to unblock an incidental edit converts a safety gate into paperwork.
infrastructure, not a capabilityThe triage above asks whether your capability actually needs the header to
change. Sometimes the header change is the whole change: a public header can
be unbuildable on a platform and the fix has to land in it. simd_buffer.hpp
wrapped C11 aligned_alloc, which Bionic does not provide below Android API 26,
so every Android build of anything including it failed to compile.
That still trips the frozen-baseline gate, and it still resolves by declaring
the header in REVIEWED_HEADERS — but declare it honestly:
"disposition": "infrastructure" with "capability_keys": []. The edit
changed no consumer contract, so there is no capability to claim. Inventing
one to look tidier would assert a binding nothing offers.contract-history.json
snapshot per the guidance above and keep the three-file change.Confirm the same post-conditions as any classification: the baseline file has no
diff, its entry count is unchanged, --check reports fresh, and the surface's
legacy_unreviewed count dropped by exactly one.
Measure --write's exit status unpiped. On a frozen header it prints
INVALID: public header fingerprint changed and exits 1, exactly as documented
above — but --write | tail reports tail's status instead, which reads as a
silent success and invites the wrong conclusion that the generator accepted the
edit.
test_agent_capability_rederive.py and test_agent_capability_manifest.py both
rewrite tracked files in place — including tools/scripts/agent_capability_manifest.py
itself — and restore them in a finally. Between the write and the restore the
repository is dirty by construction, and under -j contention that window is
the test's entire runtime rather than the fraction of a second it looks like.
That is why these tests hold RESOURCE_LOCK agent-capability-manifest-source.
The lock is not about the two tests colliding with each other; it is about
excluding anything that observes repository-wide state, which today includes
a CLI contract asserting the source tree is pristine. When you add a test here
that writes a tracked file, put it in that lock group. When you add one elsewhere
that reads global tree state, put it there too.
Declare PROCESSORS on anything in this family. The manifest self-test already
does; a sibling that omits it lets ctest schedule eight more tests alongside it,
which is how a test that runs in under twenty seconds locally reaches the
inherited 120s timeout on a loaded runner. Declaring the real cost is scheduling
accuracy, not a loosened budget.
__pycache__FROZEN_LEGACY_COUNT = 338 to 337, and one 64-hex digest to another, both
leave the source file byte size UNCHANGED. CPython invalidates bytecode on
(source mtime, source size), so a .pyc written moments earlier can survive an
edit that changed neither, and the interpreter keeps executing the OLD constant.
The symptom is a contradiction that looks impossible: python3 tools/scripts/agent_capability_manifest.py --check prints fresh, while the
identical agent-capability-manifest-check ctest fails against the pre-edit
value. The two ran DIFFERENT interpreters (ctest uses the CMake-resolved
Python3_EXECUTABLE, often a specific python3.N), each with its own
cpython-3N.pyc, and only one cache was stale.
Before believing either result, reproduce with the interpreter the test actually uses:
PY=$(grep -m1 "Python3_EXECUTABLE:" build/CMakeCache.txt | cut -d= -f2)
"$PY" tools/scripts/agent_capability_manifest.py --check
and clear the caches when a constant edit did not change file size:
find tools -name __pycache__ -type d -exec rm -rf {} +
If --check reports these on a clean checkout whose diff touches no capability
file, stop and look at the base before touching anything:
agent-capabilities: STALE: capability history is not append-only relative to the protected base
agent-capabilities: STALE: manifest changed without a manifest_revision increase
agent-capabilities: STALE: public surface changed without an inventory_version increase
The check is not wrong — it is answering correctly against the wrong reference. The protected base must be an ancestor of the commit under validation, or "append-only relative to the base" is ill-posed: measured against a tip that carries commits your branch does not have, every correct branch looks non-append-only.
Naming a moving ref makes that routine. The build hosts run ~134 worktrees off
one shared .git, so a git fetch in any sibling advances origin/main
for all of them — mid-validation included, with your session issuing no fetch.
_resolve_local_base therefore checks ancestry and steps back to the merge-base
when the ref has moved past you. The merge-base does not move when the tip
advances, which is what makes the verdict reproducible.
Two consequences worth knowing:
.git.To pin the base explicitly — for a bisect, or to reproduce a CI verdict exactly:
PULP_AGENT_CAPABILITY_BASE_REF=<sha> python3 tools/scripts/agent_capability_manifest.py --check
That path is deliberately literal: an explicit ref is used as given, without the ancestry fallback.
gates.sh runs one check the capability transaction never mentions: the exposure ledgerThe inverse of the section below is also true and catches people going the other
way. While an agent_capability_catalog_*.py file is an exclusively owned path
in docs/status/sequencer-exposure, adding a capability(...) block to it makes
gates.sh fail with
transition: sequencer-owned changed path is not covered by an added or
materially changed pending row: tools/scripts/agent_capability_catalog_<domain>.py
Nothing in agent_capability_manifest.py --check predicts this — it reports
fresh while the push is still blocked — and the four-edit checklist below is
silent about it because the fifth edit lives in a different gate entirely. The
fix is a new pending row under docs/status/sequencer-exposure/rows/
owning that catalog file and its own row file; an existing row that already owns
the path does not satisfy the transition rule. A capability published on an
installed header fills installed_sdk and design_time_agent_manifest as
exposed and the three timeline surfaces as not_applicable; measure
installed_sdk against the per-subsystem install(DIRECTORY …) loop in
tools/cmake/PulpInstallRules.cmake rather than assuming it.
Do not generalise that to "every catalog is watched" — most are not, and the
owner count is what decides. _exclusively_owned_paths in
tools/scripts/sequencer_exposure_check.py watches a path only while exactly
one row or tombstone declares it, in both the base and the resulting ledger
state. Zero owners is not watched, and two or more are shared by construction and
are not watched either. So the recorded fix is self-limiting: the row added to
satisfy the gate is another owner, and once a catalog has two, the next
capability(...) block added to it passes this gate silently. Measured on this
tree, only agent_capability_catalog_performance.py is watched at all;
agent_capability_catalog_timing.py already carried two owners before this row
existed, and foundations/signal carry none.
Count the owners of the file you are about to touch, and count them through the
checker's own loader. A glob over docs/status/sequencer-exposure/rows/*.json
gives the wrong answer three ways: one file can carry several rows, released rows
do not all live there, and tombstones declare owned_paths too. _declared_owners
is the authority:
python3 - <<'EOF'
import pathlib, sys
sys.path.insert(0, "tools/scripts")
import sequencer_exposure_check as check
root = pathlib.Path(".").resolve()
ledger = check.load_ledger_from_worktree(root)
owners = check._declared_owners(ledger[0] if isinstance(ledger, tuple) else ledger)
for path in sorted(p for p in owners if "agent_capability_catalog_" in p):
print(len(owners[path]), path, sorted(owners[path]))
EOF
A catalog printing 1 is watched by this gate; 0, or 2 and up, is not.
gates.sh does NOT run the capability check — adding a public header passes pre-push and fails in CIThe pre-push gates cover skill-sync, version-bump, compat, deps and friends. They do not
run agent_capability_manifest.py --check. So a change that adds a header under a covered
root, or edits an existing one, sails through gates.sh: all gates pass and then fails CI on
agent-capability-manifest-check / -selftest.
Adding one new DSP header produces two failures, not one:
unclassified public header: pulp/signal/<new>.hpp
public header fingerprint changed: pulp/signal/signal.hpp <- the umbrella include
The umbrella one is the easiest to miss: adding #include <pulp/signal/foo.hpp> to
signal.hpp changes that header's bytes too.
Run python3 tools/scripts/agent_capability_manifest.py --check yourself before pushing any
change under core/*/include/. Gates passing is not evidence here.
Classify honestly: a reusable bounded surface that is not an advertised generator claim takes
infrastructure with empty capability_keys and a durable rationale. Do not manufacture a
capability row to clear the gate; a capability row is a consumer contract with typed bindings
and operational probes.
A header in the frozen legacy baseline does NOT require unfreezing anything.
tools/agent-capabilities/legacy-unreviewed-baseline.json snapshots the public headers that
predate classification, guarded by FROZEN_LEGACY_DIGEST and FROZEN_LEGACY_COUNT in
tools/scripts/agent_capability_surface.py. Changing one fails with public header fingerprint changed, which reads like it demands editing those pinned constants — it does not, and editing
them to make one PR pass would be removing a deliberate guard. Declaring the header in
agent_capability_registry.py, as a capability or as an infrastructure disposition, satisfies
the fingerprint check on its own; the baseline file, its entry count, and its digest all stay
untouched. Confirm afterwards that the baseline entry count is unchanged and the manifest
self-tests still pass.
Ask the prior question first: does your capability actually need that header to change? Both this section and the recorded precedents jump straight to how to classify a frozen header, which quietly assumes the edit to it is load-bearing. Often it is not. A cell that adds a new processor and, while it is in there, refactors three existing processors onto a shared kernel will trip this gate on headers its capability never touches — and the whole gate disappears if the refactor is dropped.
So triage the edit before classifying the header:
frequency_response.hpp templated over SampleType so the _64 variants could compute a
response without narrowing: reverting it would have shipped those variants with no response
inspection, so there was no smaller correct slice.The asymmetry is what makes this worth doing in that order. Reverting an incidental edit costs
one git checkout origin/main -- <header>. Declaring a header is permanent: it converts an
unreviewed legacy header into a reviewed contract the repo then owns, decided as a side effect
of a cleanup rather than on its own merits. If the dedup is worth having, it is worth its own
change, where the classification is the subject of review instead of collateral.
When you do revert, prune whatever the reverted call sites were the only users of. A shared helper introduced for three call sites that no longer exist is dead public surface, and it enlarges the very fingerprint you are trying to keep small.
The compatibility vocabulary intentionally publishes a bounded method list per type. A conditional platform implementation overload can therefore crowd out a portable API even though the public header still contains both. When adding such an overload, exclude its platform-only signature from the generator-facing compatibility projection and add a regression that asserts both sides: the portable method remains advertised and the conditional implementation signature is absent. Do not raise the global method cap to hide this local classification error.
SURFACE_INVENTORY_VERSION is a shared ledger too — pick it by survey, not by incrementEvery branch that moves a public header has to raise it, so concurrent branches contend for the same integer. The obvious move — read main's value and add one — is wrong whenever anyone else is mid-flight, and it fails in the quietest possible way: two branches that both write the same number produce identical text, so git finds nothing to conflict on and both auto-merge clean. The collision surfaces later, at the merge commit, as an inventory version that did not actually increase over the branch that landed first.
Choose max(all live branches) + 1, not main + 1:
git for-each-ref --format='%(refname)' refs/remotes/origin \
| grep -v -- '--help\|/HEAD$' > /tmp/refs.txt
xargs -n 300 sh -c \
'git grep -h "^SURFACE_INVENTORY_VERSION" "$@" -- tools/scripts/agent_capability_manifest.py' _ \
< /tmp/refs.txt | grep -o '[0-9][0-9]*$' | sort -n | uniq -c | tail
Two details that are load-bearing, because getting either wrong returns an empty result rather than an error — and an empty survey reads exactly like "nobody holds a number", which is the answer that makes you collide:
--. Anything after --
is a pathspec, so git grep PATTERN -- path ref1 ref2 searches no revisions
and matches nothing.xargs has no -a. xargs -a file … aborts with invalid option;
redirect the file in with < instead. BSD and GNU differ here and the BSD
failure is easy to miss inside a pipeline.So pair the survey with a control that must return non-zero — git grep the
same constant on origin/main alone, which is known to carry it. If the control
is silent the instrument is broken and the survey proved nothing. A gap in the
observed numbers is not an invitation to fill it: prefer one above the maximum,
since a gap usually means that branch already landed or was deleted.
test_signal_no_exceptions.cpp is a shared ledger — three hazards, not twoNearly every signal capability appends to it, and it assigns a unique non-zero exit code per checked capability so each failure is identifiable from the status alone. Cells prepared in parallel therefore collide in it constantly. Two hazards are well known; the third is not, and it defeats the checks for the other two.
Duplicate exit codes. Two cells both append above the same remembered maximum and take the same numbers. Compiles fine, both pass in isolation, and two distinct failures become indistinguishable. Only a duplicate-code scan finds it.
Duplicate capability blocks. Rebuilding the file from main's canonical version and
appending your block is the right resolution, but a cell with more than one commit touching
the file re-adds what the resolution already appended. This one is usually loud
(redefinition of ...).
Two statements sharing a line. A resolution can leave a declaration on the same line as
the preceding return:
if (!(formants.configure(recipe) == FormantConfigureStatus::configured))
return 29; pulp::signal::ParallelDynamicsMixer parallel_dynamics;
if (!parallel_dynamics.prepare(8u, 16u))
return 31;
This compiles, and it is reachable — the if body is only the return, so the declaration
still executes — so no test can see it and neither the duplicate-code nor the
duplicate-block check fires. Both of those are line-oriented, and two statements on one
line is precisely the shape that slips past them. Observed on a published PR in 2026-08.
It is not a correctness bug today, which is why it survives review; it is a latent one. The
same shape with a guarded early return (if (cond) return N; Type x; inside a block that
can be taken) silently skips the declaration and every check after it, and the capability's
proof quietly stops running while the binary still exits 0.
So do not hand-resolve this file. Rebuild it from git show origin/main: and re-append your
block programmatically, which removes all three by construction — that is how the line-sharing
defect above was removed, as a side effect of the correct procedure rather than by spotting it.
Then verify all three: no non-zero code repeats, each capability block appears exactly once, and
no line carries two statements.
Extract codes from every return form, including ternaries (return c ? 0 : N;) — a naive
grep -oE "return [0-9]+;" misses those and manufactures phantom collisions.
For A3 v2 terminal acceptance, never treat receipt fields as publication or trace proof. The verifier must derive protected main, the canonical receipt blob, required check identities/results, and artifact digests live, then replay the pinned analyzer over the exact trace bytes.
Editing the bytes of a header carried in
tools/agent-capabilities/legacy-unreviewed-baseline.json fails the check with
public header fingerprint changed. The baseline is pinned twice over —
FROZEN_LEGACY_COUNT and FROZEN_LEGACY_DIGEST in
agent_capability_surface.py — so editing that file to match is the laundering
the pin exists to prevent, and it fails anyway.
The supported move is to graduate the header: add an entry to
REVIEWED_HEADERS in agent_capability_registry.py with the header's new
fingerprint, a disposition (infrastructure when it binds no capability of
its own), and a rationale. build_surface_document consults reviewed before
baseline_entries, so the baseline row is simply never reached — leave that file
byte-identical. The frozen count stays 337 and its digest stays valid.
SURFACE_INVENTORY_VERSION must increase relative to
docs/status/agent-capability-surface.json as it currently sits in the working
tree — which your own previous --write already moved. So a second round of
source edits (a format_changed.sh reflow is enough, since it changes the
header's bytes and therefore its fingerprint) makes --write exit 1 with
public surface changed without an inventory_version increase, even though you
already bumped. Bumping again burns a second published identity for one change.
Reset the snapshot to the base and write once instead:
git checkout origin/main -- docs/status/agent-capability-surface.json
python3 tools/scripts/agent_capability_manifest.py --write
Corollary: run format_changed.sh before deriving the fingerprint, or
re-derive after it. A fingerprint pasted from a pre-format read is stale.
Identical bumps on two branches merge cleanly and silently reuse one published
identity, so incrementing main's value is not enough. Survey the remote branches
first — and note that in zsh a "$ref:tools/..." expansion applies the :t
history modifier and silently mangles the path, so the survey loop returns
nothing while looking like a clean negative. Drive it from Python, or verify the
loop against a ref you know carries the constant:
python3 - <<'EOF'
import subprocess, re
refs = subprocess.run(["git","for-each-ref","--sort=-committerdate",
"--format=%(refname)","refs/remotes/origin","--count=250"],
capture_output=True, text=True).stdout.split()
seen = {}
for ref in refs:
r = subprocess.run(["git","show",f"{ref}:tools/scripts/agent_capability_manifest.py"],
capture_output=True, text=True)
m = re.search(r"^SURFACE_INVENTORY_VERSION\s*=\s*(\d+)", r.stdout, re.M) if not r.returncode else None
if m: seen[ref] = int(m.group(1))
assert seen, "instrument dead - no branch yielded the constant"
print("branches read:", len(seen), "max:", max(seen.values()))
EOF
PulpInstallRules.cmake fires this gate for reasons that have nothing to do with capabilitiesThe skill-path map ties this skill to tools/cmake/PulpInstallRules.cmake,
which is right — that file decides what reaches the SDK, and a new public
header arriving there is squarely this skill's business.
But the same file also carries the SDK's non-header payload: CMake modules, plist templates, catalogs. Adding a file there because an app-bundling feature needs to ship a template trips this gate with nothing to classify.
Both outcomes are legitimate; say which one you are in rather than reaching for the bypass trailer by reflex:
PUBLIC_ROOTS note above before trusting a green --check.Skill-Update: skip.
A note costs the same as the trailer and leaves the next person something to
read.PulpInstallRules.cmake installs more than public API. The macOS ObjC cluster
under core/view/platform/mac is shipped as source, because ObjC class names
are process-global and every consumer binary has to recompile those translation
units with its own class-name suffix. A .h arriving in that install block is
a private header a consumer compiles against, not a header anyone includes.
So adding one fires this gate and there is nothing to classify: the path is
outside all six PUBLIC_ROOTS, no manifest key gains or loses a digest, and
agent_capability_manifest.py --check is green before and after. Run it anyway
and say so, rather than assuming it from the path.
The thing that does need care in that block is unrelated to capabilities and
easy to miss, so it is written down here because this is the skill an editor of
PulpInstallRules.cmake is sent to. A macOS ObjC translation unit is named in
three hand-maintained lists: core/view/CMakeLists.txt, this file, and
_pulp_view_objc_srcs in PulpUtils.cmake. Omitting the install entry makes
the EXISTS probe in _pulp_apply_view_mac_objc_suffix() fail, and that
return()s, dropping the per-binary suffix for every macOS ObjC class rather
than the new one. The binary still builds and still runs; it collides only when
a second Pulp plug-in is loaded beside it, in somebody else's host. The
mac-objc-source-list-guard ctest and its selftest exist to catch that.
The private headers those .mm files quote-include are in the same install
list, and the guard checks them too: a consumer compiles the cluster as one
generated per-binary translation unit that includes each .mm by absolute
path, so a header missing from the install breaks the consumer's compile
rather than its class names. plugin_view_host_mac_view.h (the shared
PulpPluginView / PulpGpuPluginView interfaces) and mac_text_input_ranges.h
are the ones a reader is least likely to expect there.
--check says nothing about a module outside PUBLIC_ROOTSPUBLIC_ROOTS in tools/scripts/agent_capability_surface.py lists exactly
seven domains: audio, midi, music, playback, sequence, signal,
timebase. Headers anywhere else are not scanned, not classified, and not
fingerprinted.
That matters most at the moment it is least visible. Adding a new optional
module under core/ and exporting it — appending the target to
PULP_SDK_TARGETS and the directory to _pulp_sdk_header_subsystems in
tools/cmake/PulpInstallRules.cmake — ships its public headers in the SDK. Then
agent_capability_manifest.py --check prints fresh; N keys and M public headers checked and exits 0, which reads like the new headers were reviewed.
They were not looked at.
Control the reading before trusting it:
python3 -c "import re; t=open('tools/scripts/agent_capability_surface.py').read(); \
print(re.findall(r'\"source\": \"([^\"]+)\"', t))"
If the new module's include root is absent, the manifest has no opinion about it, and the header count staying put is the expected result rather than evidence of coverage.
The PulpInstallRules.cmake edit is what fires the skill-sync gate for this
skill, and that is the right moment to make the call deliberately: does the new
domain belong in PUBLIC_ROOTS? A DSP or generator-facing surface does. An
authoring, packaging, or container surface — where "capability" would mean a
consumer contract with typed bindings and operational probes that do not exist —
does not, and adding it would mean manufacturing rows to describe headers no
generator claims. Record which way you went; silence here looks identical to
having never asked.
An unpublished domain does not read as "unknown" to the agents that consume this
manifest. It reads as "Pulp does not have this", because a capability search only
sees keys that exist. agent-capabilities.json carried 106 keys across signal,
midi, timebase, audio, music and sequence and zero timeline.* rows,
so two independent passes searching for a groove projector both found
timebase.groove-kernel — the strict non-reordering realtime subset — and never
saw timeline::GrooveTemplate, the canonical authored model the kernel's own
header points back to. One of them concluded a planned slice was impossible.
The tell is that nothing failed. Coverage was already partial with
absence_semantics: unknown, every gate was green, and the manifest was
internally consistent the whole time. A missing domain has no negative control of
its own, so when weighing whether a surface deserves a row, ask what a consumer
would search for and what it would conclude from finding nothing — not whether
anything currently complains.
When two rows are near-neighbours that differ in a load-bearing way, say so in
state_model, which the digest covers, rather than only in summary, which it
does not. timeline.groove-template names timebase.groove-kernel and states
that the authored table half may reorder events while the kernel refuses to, so
whichever one a search reaches first leads to the other.
PUBLIC_ROOTSWidening PUBLIC_ROOTS admits a whole domain to the header ledger and obliges a
reviewed disposition for every header in it. Publishing a capability does not
require that, and conflating the two turns a five-row change into a thirty-header
classification pass. build_surface() resolves a binding include that
discover_headers() did not inventory through core/*/include/<include>,
verifies its declared fingerprint, and requires only a REVIEWED_MINIMAL_TARGETS
owner. pulp/host/* has published this way for a long time.
Take that path when the generator-facing part of a domain is a minority of its
headers — a document/authoring subsystem whose musical-context types a generator
reads, with editing, persistence, schema and interchange headers around them that
no generator consumes. The domain then reports not_inventoried with a nonzero
capability count, which is the honest reading: these specific rows are reviewed
and the rest is unknown. Widening the root instead would demand a claim about
every neighbouring header that nobody measured.
Two edits are easy to miss on this path because the first one masks the second:
DOMAINS in agent_capability_manifest.py. Until the domain is listed,
the row is rejected as an unknown domain.domain enum in docs/status/agent-capabilities.schema.json. This is
a different file from the surface schema the four-edit list above names, and
it fails only after DOMAINS already accepts the row —
$.capabilities[N].domain: 'x' is not one of [...], reported by the schema
validator rather than by the registry, so it reads like a regenerated-artifact
problem rather than a missing enum member.--check says fresh while the history is missing your keyagent_capability_rederive.py resets the generated artifacts to the protected
base before regenerating. Immediately afterwards contract-history.json can hold
a snapshot that predates the new key while agent_capability_manifest.py --check
reports fresh and exits 0 — because --check does not require the history to
contain the current contract. Measured: after a rederive a grep for
timeline.groove-template returned 0 while the control timebase.groove-kernel
returned 62; a following --write moved them to 1 and 63, and --check said
fresh in both states.
So do not read --check as proof the history recorded anything. Grep the history
for your own key alongside a key you know is already there, and compare both
counts. Then check entries: running --write after a rederive that already
wrote appends a second snapshot (61 → 63, ~35k lines for a change that needs
~18k). Restore the file from the protected base and run --write exactly once so
the diff carries one snapshot.
PUBLIC_ROOTS is four more edits, and each hides the nextThe four coordinated edits above cover a new header inside a domain that is already scanned. Admitting a whole new domain is a second, disjoint set, and the tooling reveals them strictly one at a time — satisfying one produces an error that looks unrelated to the one before it:
PUBLIC_ROOTS in agent_capability_surface.py. Until the root is
declared, discover_headers() returns nothing for it, so the domain reads as
fully reviewed because it is entirely invisible.docs/status/agent-capability-surface.schema.json —
in three separate places (the review-root domain, the inventory's
propertyNames, and the frozen-entry domain). Miss one and a row that the
surface script just legitimately produced is rejected as schema-invalid.REVIEWED_MINIMAL_TARGETS in agent_capability_registry.py. The failure
is include has no covered public target owner, which names the binding, not
the missing map row. The value is the CMake export name
(Pulp::playback), which PulpInstallRules.cmake derives from the target
(pulp-playback) — not the target name itself.test/cmake/quality_tests.cmake. The
generated link probe for pulp-test-agent-capability-compile cannot resolve a
symbol from a subsystem the target does not link, so the probe fails at link
time with no mention of capabilities at all.Two preconditions are worth measuring before starting, because assuming either
one wastes the whole pass. The domain's headers must already install — the check
is the per-subsystem install(DIRECTORY …) loop over
_pulp_sdk_header_subsystems in PulpInstallRules.cmake, not the presence of
an install(TARGETS …) line. And the target must already be in
PULP_SDK_TARGETS; if it is not, adding it is itself public surface (see below)
and belongs in its own slice.
Partial coverage is the expected end state, and it must be spelled. Classify the
headers that are not part of the advertised closure as infrastructure with
empty capability_keys, never as unsupported_capability: an absent key means
unknown, and unsupported_capability asserts a fact about the header that
nobody measured.
tools/cmake/PulpInstallRules.cmake holds PULP_SDK_TARGETS, the list
cmake --install exports. If you split a target that appears in that list into
an umbrella plus its halves, every half must be added to the list too, not
just kept behind the umbrella name.
The trap is that the umbrella still installs fine on its own. What breaks is
the consumer: the exported umbrella's INTERFACE_LINK_LIBRARIES names the
halves, so a downstream find_package(Pulp) resolves a target whose interface
references targets the export set never defined, and fails there rather than at
install time. The symptom appears in someone else's build, one step removed
from the change that caused it.
pulp-format is the worked example: it exports as pulp-format,
pulp-format-core and pulp-format-view. Note also that the umbrella must
stay a real STATIC library rather than INTERFACE, because the export set
expects an archive artifact.
PULP_SDK_TARGETS is public surfacePULP_SDK_TARGETS in tools/cmake/PulpInstallRules.cmake is the export set, so
appending a target publishes a new Pulp::<name> that outside projects can name
in target_link_libraries and find_package(Pulp COMPONENTS ...). It carries
the ordinary compatibility weight even when the target is INTERFACE-only and
ships no archive.
Anything already in an exported target's link interface MUST be in that set: CMake refuses to export a target whose interface names one that is not. So an INTERFACE target introduced to narrow another target's link line is not optional to export, it is required by the export that motivated it.
A target defined under core/<x>/ but not owning a core/ directory of its own
draws a module '<x>': CMake links pulp-<name> but modules.yaml doesn't list it
warning from tools/check-docs.sh. That warning is correct and unfixable from
modules.yaml, whose entries are validated against core/<name>/ existing;
adding a row would convert a warning into a hard failure. pulp-tracing,
pulp-perfetto and pulp-cpp sit in the same position.
quality_tests.cmake must declare PROCESSORSThe capability tests (agent-capability-manifest-check,
agent-capability-manifest-selftest, agent-capability-rederive-selftest)
share test/cmake/quality_tests.cmake with much heavier GPU and role-producer
selftests. CI runs that suite with ctest -j8 --timeout 120, and ctest's
scheduler charges a test one slot unless PROCESSORS says otherwise.
A test that spawns a subprocess tree while declaring the default single slot is
therefore co-scheduled with seven other tests that may do the same. Each then
inflates the others, and the ones nearest their budget time out — on a loaded
host, tests that pass comfortably in isolation fail together in a cohort, on
unrelated PRs, in varying subsets. The failure looks like flakiness or like a
break on main; it is neither.
When you register a test here:
PROCESSORS 8 if it forks, builds, or drives subprocesses.
agent-capability-manifest-selftest already does.TIMEOUT when the work genuinely exceeds the suite
default. build.yml sets --timeout 120 specifically so one hung entry
cannot burn the workflow timeout, and its comment directs long tests to set
their own. An explicit TIMEOUT relaxes no assertion — every negative
control still runs and still must refuse.agent-capability-rederive-selftest takes 76.65 s and
gpu-first-visible-role-producers-selftest 49.07 s. The first has roughly
1.6x headroom and the second 2.4x, and both have timed out in the same
cohort — so a slot declaration without a budget still leaves the tighter one
failing under load.git show
hashing that costs seconds over the real tree ran against a two-file
synthetic fixture and accounted for under 1% of runtime, while the true cost
was the count of sealed-build invocations.RESOURCE_LOCK agent-capability-manifest-source serializes the three capability
tests against each other, but it constrains nothing against tests holding a
different lock — it is a correctness guard for the manifest rewrite, not a
concurrency budget.
Measure on the slowest runner the gate can land on, not the one you have.
The required macos gate places on either Mac Studio, and M5 runs the same
suite roughly 1.5x slower than M3. A selftest measured at ~62 s on an idle M3 —
gpu-recipe-catalog-selftest, which builds throwaway clones per equivalence
class — is already over the 120 s default once that factor and a loaded host are
applied, even though the local number looks like comfortable headroom. Scale the
local measurement before deciding a test needs no explicit TIMEOUT.
inspect/ control header is not a design-time capability rowtools/cmake/PulpInstallRules.cmake installs the capability-control executor
headers (control_state_write_executor.hpp,
control_timeline_document_session_executor.hpp, and their siblings) into the
same SDK that carries the design-time agent-capability contracts. Sharing an
install list is not sharing a registry, and the resemblance is the trap: the
headers declare typed request/outcome structs and a resolver, which reads like a
binding, so the reflex is to add a manifest row "for consistency".
Do not. Runtime operation metadata — operations, grants, instances, receipts —
is broker authority and stays outside the design-time manifest by construction.
A control operation is declared once in inspect/src/control_manifest.cpp and
its capability once in capability_definitions.inc; the CLI and MCP surfaces
then project it from control_operation_registry(). Nothing in that path reads
the agent-capabilities manifest, so a row added there would advertise a contract
no consumer resolves and no gate re-derives.
The practical consequence is a disposition, not a code change: a sequencer
exposure row covering a live control operation records
design_time_agent_manifest: not_applicable with that boundary as its
rationale, and never gap — gap claims someone owes the row, and nobody does.
What the install entry does buy is honesty on a different surface. Adding the
header to the install list is exactly what lets the same exposure row claim
installed_sdk: exposed, because an embedding host then links the typed source
seam rather than re-declaring it. Omit the install entry and that claim is
false, while the manifest row would still have been wrong.
midi.humanize 1.1 advertises the compatible future-attack spec update method
and supports a nonnegative timing floor. Its operational binding constructs a
kernel and invokes the update with a bounded spec. This design-time kernel
registration does not advertise placed-device parameter operations or grants;
the event-humaniser exposure ledger keeps those product-control gaps explicit.
A sample-region capability describes a declared graph/control contract. Region proof and runtime admission are the authority; capability rows and inspector examples are projections. Keep unsupported latency, state-size, history, and format projections explicit until their contracts and independent proofs exist. Do not use vestigial flags or create a second DSP registry.
A graph-owned custom-node diagnostics query can change a catalog-bound
signal_graph_runtime.hpp fingerprint without changing sample-region authoring
semantics. Refresh every existing binding to that header, but do not advertise
instance counters, availability, generation handles or provider reports as new
design-time capabilities. The sibling diagnostics descriptor is a runtime
inspection contract. Its new header follows the host directory install/Doxygen
parity rules; do not expand the deliberately closed sample-region header list
merely to force runtime diagnostics into the design-time catalog.
Adding a member to a fingerprinted public header (here Processor::editor_prewarm()
in pulp/format/processor.hpp) makes the pre-push agent-capability gate report
public header fingerprint mismatch, and agent_capability_rederive.py refuses
to derive counters until it is fixed. The hash lives in the owning catalog
(tools/scripts/agent_capability_catalog_*.py, header_fingerprint=) and, when a
pending sequencer-exposure row pins that catalog, in that row's evidence needles
too (git grep the old hash). Replace both with shasum -a 256 of the header,
then re-run the rederive with PULP_AGENT_CAPABILITY_BASE_REF=$(git rev-parse origin/main); it reports whether the counters themselves need to move (often
they do not), and agent_capability_manifest.py --check plus
sequencer_exposure_check.py --base origin/main confirm both ledgers.
--writeupdated_history_entries appends the manifest that was on disk BEFORE the
write (and only when it differs from the last entry), not the manifest being
written. So the first --write of a contract change re-records the base state,
appends nothing, and still reports writing contract-history.json; --check
says fresh. Moving the MIDI routers to major 2 showed it: from a base history
the first --write left 84 entries ending at major 1, and a second --write
produced 85 ending at major 2. Until the tool records the current manifest
itself, confirm entries[-1].manifest.capabilities[<key>].contract_version
shows the new version and that the file differs from the base by exactly one
entry before shipping.
pulp/host/signal_graph_authoring.hpp is a reviewed infrastructure header for
GraphAuthoringReceipt and GraphAuthoringReceiptStatus. Keep it in the
reviewed host tuple and minimal-target registry with an empty capability-key
list; it describes lineage validation vocabulary and does not advertise a DSP
capability. Refresh its byte fingerprint and the surface inventory when the
contract changes.
まだレビューはありません。使ってみた感想をお寄せください。
概要と使いどころ
Optional AAX support for Pulp, including developer-supplied Avid SDK setup, CMake enablement, DigiShell/AAX Validator workflows, and local AAX builds on macOS or Windows.
日本語の概要は準備中です。原文の説明を表示しています。
Configure, implement, and test Pulp's optional desktop Ableton Link tempo-sync adapter while preserving the developer-supplied SDK, licensing, realtime, latency-compensation, and no-install boundaries.
日本語の概要は準備中です。原文の説明を表示しています。
Android platform development for Pulp — NDK cross-compilation, Oboe audio, Dawn/Skia GPU rendering, JNI bridge, touch interaction, emulator workflows, and end-to-end smoke validation. Covers build, deploy, debug, and the gotchas discovered during bringup.
日本語の概要は準備中です。原文の説明を表示しています。
Optional ARA support for Pulp, including developer-supplied ARA SDK setup, CMake enablement, adapter companion APIs, validation, and ARA-aware plugin implementation guidance.
日本語の概要は準備中です。原文の説明を表示しています。
The measurement surface for ALL Pulp DSP and audio-pipeline work — read it BEFORE writing or gating DSP, not only when something already sounds wrong. Covers the C++ harness (signal generators, metrics, assertions, RenderScenario, contracts), the offline Audio Doctor (magnitude/frequency response, THD/THD+N, phase/group delay), and their Python sibling the Audio Quality Lab (tools/audio/quality-lab — null residual + alignment, LTAS log-spectral distance, spectral flux/centroid, HNR, Theil-Sen drift slope, Kaiser-sinc resampling, license-guarded corpus, regression-net ratchet). TRIGGER on AUTHORING work — "build/design an oscillator/filter/synth/effect", "add a DSP module", "what should the acceptance gate be", "how do I measure aliasing / anti-aliasing / alias floor", "null against a reference", "is this DSP correct", "choose a tolerance", "golden/regression corpus for audio", "measure drift or jitter", "A/B two renders" — AND on DEBUGGING work — "is there sound / no audio / I hear nothing", "does this filter/compressor/synth/delay produce the right signal", "prove the DSP / prove the contract", "measure the frequency response", "what's the THD / is it distorting", "what's the group delay / phase response / measured latency", "magnitude response curve", "render a test tone and assert", "audio regression", "64-frame works but 128 is silent", "sample-rate change pitch-shifted it", "describe what's in this buffer", "audio doctor", "compare before/after a DSP refactor". Reach for this BEFORE hand-rolling any FFT, null test, alias measurement, pitch tracker, or golden-render script — most of it already exists in one of the two lanes. Test/tool layer over HeadlessHost — deterministic, no audio device, no speakers. Off the realtime thread entirely.
日本語の概要は準備中です。原文の説明を表示しています。
Reproduce and debug "only happens in a DAW" audio plugin bugs (cutouts, glitches, parameter-change failures) entirely offline — headless Processor scenes for DSP bugs and a standalone AudioUnit host probe for adapter/host-interaction bugs. Use when a plugin misbehaves in Logic/Live/etc. but unit tests are green.
日本語の概要は準備中です。原文の説明を表示しています。