本文へ移動
cccskills
無料GitHub で公開

trace-analysis

The investigation harness for "why is this slow?" over a Pulp Perfetto trace (.pftrace). Runs the hypothesis→query→drill-down chain-of-evidence loop autonomously and returns a plain-English root cause + evidence + a concrete fix. TRIGGER on "why is my plugin slow to open", "find the slowest frames", "why is the UI stuttering", "why is my plugin using so much CPU", "which DSP node is expensive", "the load meter looks calm but CPU is pinned", or a named `pulp trace gpu-startup|gpu-health|gpu-probe` analysis. Ships Pulp-specific hints for dsp, frame, js, gpu, and cross-platform symptoms.

インストール方法を見る

含まれるファイル(9)

  • SKILL.md80.8 KB
  • references/hints_crossplatform.md3.7 KB
  • references/hints_dsp.md6.2 KB
  • references/hints_frame.md4.7 KB
  • references/hints_gpu_audio.md5.1 KB
  • references/hints_gpu.md4.8 KB
  • references/hints_js.md2.9 KB
  • references/ui_jank_playbook.md11.7 KB
  • references/ui_jank.sql10.2 KB

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

trace-analysis — the "why is this slow?" investigation harness

Someone who has never opened a profiler types one line — "why is my plugin slow to open?" — and gets a real, plain-English answer with a suggested fix. This skill is the protocol that makes that happen: you load it, capture or accept a .pftrace, and run a disciplined investigation over it, returning a narrated root cause and a chain of evidence — never raw SQL. You are reading this because a tracing question fired (L1 explain) or an expert is driving an iterative investigation (L2).

This skill is provider-agnostic by construction: it is a plain SKILL.md read identically by Claude Code and Codex from .agents/skills/. The whole "why is this slow?" experience works under any agent because it is a shipped skill + CLI/MCP surface, not a vendor API.

Its companion is trace-sql — the SQL discipline and the Pulp trace-stdlib of named query primitives. This harness decides what to ask; trace-sql is how to ask it. Load both for an L2 session.

Attribution. The investigation methodology below — the chain-of-evidence scratchpad, the hypothesis→query→drill-down loop, the wall-time-vs-CPU-time rule, follow-the-blocker, exhaustive global verification, and p95/p99 discipline — is adapted from Google's android/skills perfetto-trace-analysis (Apache-2.0). The methodology is reused; the Android domain content (SurfaceFlinger, binder, RenderThread, ftrace, cpu governor) is not — the Pulp domain hints in references/ are authored fresh against Pulp's own seams. See NOTICE.md.


Measurement mistakes that produce confident wrong answers

Read this before quoting any number. Every row happened on one live-editor investigation (a spectrum analyzer with modulation, audio playing, pointer drawing). None of them errored. Each produced a clean, plausible result that was wrong — and several survived several rounds of "fixes" measured against them. A run that fails any Detect check is discarded, not averaged in.

MistakeWhat it producedHow to detect itRule
Trace ring wrappedA 512 MB ring over 60 s with audio (~715 MB written) silently dropped the main-thread sequence: 14 of 15 runs read as "no frames".stats: traced_buf_incremental_sequences_dropped > 0 or any wrap; file size ≈ ring size; trace_processor prints "Trace health issues: Data losses".Size PULP_TRACE_RING_KB so the whole run fits (1572864 = 1.5 GB for 60 s with audio). Run the health query on every capture, first.
No audio flowingSilence / a test fixture left the analyzer path idle; the real cost (audio through the plugin) was invisible.Data-delivery slice or counter near zero; meters flat in the screenshot.Always drive PULP_TEST_SIGNAL=noise (or sine). Do not treat "no input" as "analyzer idle" either — silence still sends analyzer frames.
Wrong window drivenA global CGEvent mouse mover ran while the app was not frontmost and dragged the user's terminal instead.Frontmost pid ≠ app pid at any point during the run.Use in-window stimulus (PULP_TEST_POINTER_DRAG='rect:…'); gate every run on frontmost pid == app pid.
App never actually loadedLaunched from a sandboxed agent shell, the GUI app never attached to the window server (no window, not registered with LaunchServices) — yet the trace still had frame slices from a non-window source.No window id; screenshot empty; paint/gpu_acquire counts ≪ frame count.Per-run validity gates: frontmost pid, a window id plus a screenshot someone looks at, and a trace containing editor paint and data-delivery activity.
Screen lockedThe display link stops: zero or near-zero editor frames, a well-formed "idle" trace.Check lock state before and after each run.Discard the run.
Degraded OS audioCoreAudio held 4,096 com.apple.AirPlayXPCHelper plug-in objects; every audio app spent 19–75 s in device init, so runs started with no window.Time a kAudioHardwarePropertyDevices query; count the system object's kAudioObjectPropertyOwnedObjects by class (thousands of one class).A human restarts AirPlayXPCHelper then coreaudiod (-9), or reboots; let the load settle before measuring again.
Harness measures itselfScenario/menu-step harness snapshots added 100–800 ms stalls that read as app jank.Stalls line up with harness actions; any snapshot/eval the script runs lands in-app.Nothing the measurement script does may run inside the app during the capture window.
Shared-host loadLoad average 7 vs 30–130 moved frame times more than most fixes did.Record load average at the start and end of every run.Interleave old/new runs (A, B, A, B, …) in one session; never compare against a number from another session.
Automation instead of the gestureSetting host params (param_set) to simulate zoom/shape changes exercised the host-automation projection path (22–127 ms per change), not the user's wheel/click path.The slices in the window are host-sync/projection, not dom_event_*.Measure both paths, labelled separately; the user's gesture is the one that decides feel.
Averages hide the feelavg 16.7 ms with max 145 ms still felt sluggish.—Report p50/p95/max and counts over 50 / 100 ms per phase: mid-stroke, the 2 s after release, steady modulation, zoom.
Frame rate ≠ content smoothnessA steady 60 fps drew a spectrum that updated ~23 Hz (the analysis hop).—Measure content cadence separately: deliveries/s, max gap, gaps > 150 ms, resets.
Tracing build ≠ shipping buildTracing wraps every JS→native call in a slice, inflating absolute costs.—Compare deltas between tracing builds; judge feel on a release build.
Tooling returned a silent wrong answer/usr/bin/grep is ugrep (backreferences error out → every run marked INVALID); a zsh rm with a no-match glob aborted an && chain, so a batch "finished" instantly with RC=1; a pipeline's exit status is its last command's.A batch that finishes too fast; every run failing the same way; a zero with no control.Pair every negative finding with a positive control that must be non-zero on the same instrument (CLAUDE.md, "Pair every NEGATIVE finding…").

The copy-pasteable capture → validity → per-phase → worst-frame workflow that applies these rules is references/ui_jank_playbook.md, with its SQL in references/ui_jank.sql.


The tiers (where this skill sits)

TierWhoEntryThis skill's role
L0novice, no agentpulp trace gpu-startup|gpu-health|gpu-probe --trace FILEreturn the same bounded typed analysis used by MCP
L1novice, one-shotA .pftrace plus a questionrun this protocol autonomously, return narrated root cause + evidence + fix
L2expert, iterativepulp trace query "<sql>" --trace FILE + this skill + trace-sql loadeddrive the full loop by hand on hard/multi-bottleneck cases

The three GPU questions are closed, checked-in PerfettoSQL analyses; they never accept raw SQL or select a live target. Other planned named presets and live Trace.query / Trace.explain remain unavailable. L1 runs the broader workflow over real offline queries.


Capture (if you don't already have a .pftrace)

Getting a build that can trace at all

Tracing is off by default, so an ordinary build emits nothing no matter how pulp trace start is invoked. Do not reach for cmake -DPULP_TRACING=ON and do not copy an SDK prefix under a -trace name — a name is not evidence, and a hand-copied prefix inherits the release provenance of whatever it was copied from, which is how a directory came to claim tracing while containing zero Perfetto.

pulp build --trace                                 # in a Pulp checkout
pulp sdk install --local --profile trace --print-path   # for a standalone project

Both paths verify the result before handing it to you: the trace SDK profile refuses to publish a staged install unless PULP_TRACING is on in the cache, the runtime archive carries the tracing ship sentinel, and the exported package declares Pulp::tracing.

--trace builds into build-trace/, separate from build/, because PULP_TRACING reaches every translation unit. Expect a cold build the first time; after that the two trees rebuild independently.

Before investigating an empty trace, check the build can trace. pulp status prints a Tracing: line, and a consumer build can read PULP_HAS_TRACING from find_package(Pulp) — set from the presence of the exported target, so it is evidence rather than a claim. Silence from a build that was never configured with tracing looks identical to silence from a bug.

If a capture or query fails for an unclear reason, run the readiness check first. It reports offline trace_processor readiness without probing a legacy Inspector endpoint:

pulp trace doctor            # human report; add --json for machine output

ready_to_query:false means no trace_processor (on $PULP_TRACE_PROCESSOR, the pinned Pulp-fetched build, or $PATH) or no captured trace yet. For zero-install, run pulp trace fetch once — it downloads the pinned trace_processor_shell (Perfetto v57.2), SHA-256-verified, into $PULP_HOME. (pulp tool install trace-processor fetches the same pinned artifact via the tool registry.)

The pulp-rust-gpu-trace-analysis-integration CTest is always registered. At test time its wrapper accepts an explicit executable $PULP_TRACE_PROCESSOR or the pinned v57.2 Pulp cache. A missing processor is a visible CTest SKIP (exit 77); an invalid explicit override is a failure; a resolved processor runs the full fail-closed Cargo integration suite. Do not condition test registration on the configure-time environment or accept an arbitrary PATH version for this SDK-matched integration gate.

pulp trace start --categories render,gpu,text,js,layout   # pick the categories the question implicates
# ... reproduce (open the editor, sweep the knob, run the offline render) ...
pulp trace stop

Or accept a --trace FILE.pftrace the user hands you. Choose categories from the question: startup → render,gpu,text,js,layout; DSP cost → dsp,dsp.node (offline render); UI hitch → render,layout,canvas,text,js,gpu. Motion capture is currently available only through in-process fixture APIs.

Live start/stop use the canonical capability-control client and have no legacy Inspector fallback. If the broker cannot authorize a target, capture fails closed. When selecting an exact broker-owned live instance, pass the same --instance ID to both start and stop; omission retains fail-closed unambiguous selection.

Before trusting the capture, confirm it is not silently empty/truncated. The three named GPU questions inspect processor-reported truncation, whole-trace slice count, unfinished slices, and positive Perfetto data-loss/no-flush stats automatically. For free-form analysis, query both SELECT DISTINCT category FROM slice and SELECT name,value FROM stats WHERE severity='data_loss' AND value>0. No rows, or a missing requested category, means re-capture with a larger --ring-mb or a shorter window—not "nothing was slow."

When you already have a flushed .pftrace and no live session (the common case for a trace a user handed you), run SQL against the file directly — no inspector needed:

pulp trace query "SELECT DISTINCT category FROM slice" --trace /tmp/pulp-<ts>.pftrace

This shells out to trace_processor_shell ($PULP_TRACE_PROCESSOR → pinned Pulp-fetched build → $PATH; pulp trace doctor reports which, pulp trace fetch installs the pinned one) and returns its native table. --format json/csv and --preset apply only to the live inspector path; offline is raw SQL against the file.

To eyeball a trace in the Perfetto timeline instead of querying it, hand it to the UI (browsers block file://, so this serves it over loopback and opens the UI at it): pulp trace open /tmp/pulp-<ts>.pftrace (--no-browser prints the URL to paste; --json for agents). If the UI reports a 404 for the trace on an older CLI, that was the loopback server reading an early connection as an empty request (macOS accepted sockets inherit the listener's non-blocking mode); rerun pulp trace open or update the CLI.

For the bounded GPU path, start with one named question. gpu-startup is deliberately unverified until A3 defines a measured budget; gpu-health and gpu-probe return pass/fail. Missing categories, unfinished slices, or invalid probe evidence return unavailable rather than a misleading pass. Read cold_start_contributors separately from steady_state_contributors. Treat execution_state as CPU/wait evidence only when scheduler_evidence_available:true, which requires complete interval coverage for at least one contributor; partial/absent coverage stays unavailable. That state means the wall-clock span does not prove blocking, so use the emitted platform-harness action. Treat an unknown timing phase as neither cold nor steady.

pulp trace gpu-startup --trace /tmp/pulp.pftrace --json
pulp trace gpu-health  --trace /tmp/pulp.pftrace --json
pulp trace gpu-probe   --trace /tmp/pulp.pftrace --json

Named GPU analysis is deliberately bounded at the external-tool boundary: the trace must be at most 512 MiB, trace_processor gets one 120-second wall-clock deadline, and stdout plus stderr may retain at most 4 MiB. A deadline, output overflow, or pipe-read failure terminates the entire Unix process group or Windows Job Object, including descendants. Do not replace this runner with Command::output, child-only termination, or an unbounded wait. If a valid trace hits a limit, narrow or shorten the capture first; use L2 offline queries to split a genuinely large investigation. Treat repeated limit failures as an actionable processor/query problem, not as evidence that rendering is slow. Before launch, the named analyzer opens without following the final symlink/reparse point and without blocking on a Unix FIFO/device replacement, then verifies a regular-file handle. It copies those bytes into an exclusive private snapshot (mode 0600 on Unix) and rejects replacement or resize using handle-derived filesystem identity (device/inode on Unix, volume/file ID on Windows). The processor receives only that snapshot pathname; do not restore a stat-then-reopen flow over the caller-controlled path.


A declared category is not a populated one

trace.hpp declares ten categories — dsp, dsp.node, render, layout, canvas, text, js, gpu, state, io. Declared is not the same as emitted, and the difference reads as a finding rather than as a gap.

Some of those categories emit from nowhere. A query filtered on one of them returns no rows, which is indistinguishable from "that subsystem did no work".

Measure it; do not trust a number written here. Any count in a document goes stale the moment someone instruments something — this section previously stated a frozen count for canvas, which its own instrumentation then made wrong:

for c in dsp dsp.node render layout canvas text js gpu state io; do
  printf '%-10s %s\n' "$c" \
    "$(git grep -oE "PULP_TRACE_[A-Z_]+\(\s*\"$c\"" origin/main -- core inspect | wc -l)"
done

Run that before drawing any conclusion about a category. A zero there means the category emits from nowhere, so no trace can contain it.

The trap is expensive because the query SUCCEEDS:

-- Zero rows is NOT "this subsystem is free".
SELECT SUM(dur) FROM slice
  JOIN track ON slice.track_id = track.id
 WHERE slice.category GLOB '<category>';

Before concluding that a subsystem costs nothing, pair the question with a control on a category the sweep above shows is populated. If your target returns zero and the control returns rows, you have measured an instrumentation gap, not a fast subsystem. Say that, and stop — do not report a performance verdict about a subsystem the trace cannot see.

Where an uninstrumented subsystem's cost still shows up: inside whatever render frame scope and gpu submit/present slices contain it, aggregated. That tells you a frame was expensive; it cannot tell you which work inside it was responsible.

Per-frame cost of an animated UI, in one command

For "is this animation cheap?" -- a modulated knob, a meter, a plot an LFO moves -- the answer is per frame and per stage, split by scenario, and it has a helper. Capture with your own frame span around each frame's work (or use the host's frame / plugin_editor_frame) and a scenario span around each run, then:

tools/scripts/trace_frame_cost.py --trace cap.pftrace --frame modctl_frame \
    --group modctl_scenario --labels idle,knobs,morph,bands \
    --baseline idle --max-p95-ms 1.5 --max-layout-frames 0     # exit 1 on breach

It prints frame p50/p95/max, each stage's p95 by SELF time (so a js span wrapping canvas work is not counted twice), whole-surface repaint requests per frame and the frames that ran a layout pass; --json for a machine. It exits 2 when the frame span never fired or the ring wrapped, never reporting an empty table as a pass. The SQL is pulp_frame_stage_cost in the trace-sql stdlib.

What to read, in order:

  1. Frames that ran layout -- a value tick that writes a size or a React commit shows here first; it should be 0 for a modulated control.
  2. Frame p95 over the baseline -- the cost the animation adds.
  3. repaint/f -- whole-surface requests per frame. A frame-drive with zero bounded damage usually means the bridge's own repaint request (repaint_request → view_repaint_request) or a bounded request escalated by View::request_repaint(Rect) (render transform, filter or scroll on an ancestor -- a design-viewport scale is one). Find who asked: SELECT p.name, a.display_value, COUNT(*) FROM slice s JOIN slice p ON s.parent_id = p.id LEFT JOIN args a ON a.arg_set_id = p.arg_set_id WHERE s.name = 'view_repaint_request' GROUP BY 1, 2 ORDER BY 3 DESC -- the js_native parent's debug.fn names the bridge call (setSvgPath, ...).

Pair the trace with the headless count gate (FrameCostProbe, view-bridge skill, checklist item 8) and its negative control. Getting a traced build: pulp sdk install --local --profile trace or a downstream's own trace-SDK script; verify it by symbol count before trusting an empty capture (nm lib/libpulp-perfetto.a | grep -c " T " in the thousands, nm lib/libpulp-view-core.a | grep -ci perfetto non-zero), and note that a PULP_TRACE_SCOPE_NAMED with a category Pulp does not define (core/runtime/include/pulp/runtime/trace.hpp) only fails to COMPILE in a traced build -- a release SDK hides it.

The investigation protocol

1. Keep a chain-of-evidence scratchpad

Write down, as you go: the question → your current hypothesis → the query you ran → what it showed → the next hypothesis. Every claim in your final answer must trace back to a specific query result. "Looks slow" is not a finding; "a text span of 540 ms precedes the first frame, from query X" is.

2. Hypothesis → query → drill-down loop

Form one hypothesis at a time, query it via the trace-sql stdlib (start with the named views), read the result, refine or discard. Drill from coarse to fine: category totals (pulp_layout_vs_paint) → the fat span → the thread it ran on → its children → its args. Do not write one giant query; walk down.

3. Wall time vs CPU time — the rule that prevents wrong answers

slice.dur is wall-clock duration. A long slice may be blocked (waiting on a lock, I/O, the GPU, another thread), not computing. Before blaming a long span, check whether it was actually running: join to thread_state / inspect the thread's scheduling for that window. A 600 ms span that was Runnable/blocked 90% of the time is a waiting problem (fix the blocker), not a compute problem (optimize the code). Getting this backwards sends the fix in the wrong direction.

A userspace-only capture cannot answer this, and says so by returning nothing. thread_state and sched are populated from ftrace; a trace with only the in-process SDK's track events has zero rows in both, and zero slices carrying thread_dur. Check that before reading a join's empty result as "never blocked" — count the rows first, and if the instrument is empty, say the split is unavailable rather than reporting a wall-time number as CPU time. The fallback that does work on such a trace is descendant self-time: subtract each slice's children from its own duration and rank what is left. Self time inside a JS evaluation span, with individually trivial bridge calls beneath it, means the cost is the script, not the work it asked for.

4. Follow the blocker across threads

When step 3 says a span was blocked, follow the blocker: which thread / resource held it? The audio block waited on a mutex the UI thread took; the present waited on a GPU pass; the layout waited on a JS callback. Trace the wait to its holder on the other thread and make that the next hypothesis. The bottleneck is frequently not on the thread that looks slow.

5. Exhaustive global verification — do not stop at the first bottleneck

Finding one fat span is not the end. Verify globally: is this the dominant cost, or one of several? Sum the category totals and check the found span's share. Look for a second offender of comparable size. A fix that removes a 300 ms span from a 2.4 s open still leaves 2.1 s — say so. Report the full budget, ranked, not just the first thing you noticed.

6. p95/p99 discipline — the mean lies

For anything recurring (per-block DSP, per-frame render), the mean hides the spike. A node at 40% mean can eat 60% of the worst block. Look at max_ms next to mean_ms (the stdlib views expose both), and when the gap is ambiguous compute the tail (p95/p99) over the raw slices — see trace-sql. "The load meter said 40%; the trace said WHY" is exactly this: the average was calm, one node's tail was not.

The mean also lies in the other direction: a per-node average can be dominated by a one-time cold-start spike on the first block (the first process() call warms caches / touches fresh pages), making a genuinely cheap node look like the worst offender. Always separate the first-block outlier from steady-state cost — subtract the per-node MAX(dur) (the steady-state query in trace-sql's "Common query shapes"), or just read the flamegraph and ignore block 0. examples/trace-plugin-chain is a runnable demonstration: gain's whole-run average is ~158× its steady-state cost, and the real per-block hot node is the biquad filter, not the gain the average fingers.

7. Consult the domain hints

Match the symptom to a hints file and read it before drawing conclusions — each grounds the analysis in Pulp's real seams and names the specific traps:

SymptomHints file
xruns, per-node DSP cost, deadline miss, "meter calm but one node dominates", jitter/denormalsreferences/hints_dsp.md
"it clicks / drops out here": a glitch in a render joined to the block that rendered it, block time vs deadlinereferences/hints_dsp.md ("Block time against the deadline"), tools/audio/glitch_trace.py
dropped frames vs vsync budget, layout-vs-paint, TextShaper::prepare re-runs, dirty-rect churn, GPU-submit stallsreferences/hints_frame.md
QuickJS bridge dispatch cost, a JS callback invalidating layoutreferences/hints_js.md
Dawn submit/present stalls, Graphite record cost, per-pass GPU timereferences/hints_gpu.md
Shared-I/O GPU audio admission, terminal/delivery correlation, and quiescent recoveryreferences/hints_gpu_audio.md
a drag/scroll/hover that feels sluggish while frame medians look fine; frames stall only while the mouse moves; huge bridge-call counts over one interactiondocs/guides/interaction-cost.md
a live editor (meters, analyzer, modulation) that feels sluggish mid-stroke, after release, or during zoom; classifying the worst frame gapsreferences/ui_jank_playbook.md (+ references/ui_jank.sql)
standalone vs plugin-in-DAW vs iOS/iPadOS AUv3 vs Android/Oboe vs Simulator; sample-position args, thread naming, atrace interleavereferences/hints_crossplatform.md

GPU-audio scheduling extensions may annotate nullable batch, model, provider, and execution-prediction metadata. Read the companion hint before treating a lead or batching result as a claim: session claim_* flags opt into fail-closed missing_* checks, and absent historical schema-2 fields remain unavailable rather than zero. A negative deadline margin is valid signed slack; only NULL means that prediction metadata was not emitted.

8. Answer in plain English (L1) — never surface SQL

Return: root cause (one or two sentences), chain of evidence (numbered, each tied to what a span/query showed), and a concrete fix. Give magnitudes ("~620 ms Dawn/Graphite init, one-time" ), say whether the cost repeats, and estimate the win. See the worked narrative below.

Escalating to the user — only at a genuine priority fork

Default: investigate autonomously. Follow the blocker, gather evidence, and verify exhaustively. Resolve every factual gap from the trace itself — a category total, a thread state, a span's args are things you query, not things you ask. Never stop to ask the user something a query can answer.

Reach for AskUserQuestion only when the direction turns on the user's priorities, not on data you can gather. Genuine forks: startup splits near-evenly across font-shaping and shader-compile and only the user knows which path matters for their use case; or the fix has a fast-approximate branch and a slower-thorough branch and the tradeoff is theirs to pick. These are preference forks, not missing facts.

Form it well: put the recommended option first and label it recommended, with a terse pro/con per option. Never use it as a progress checkpoint ("should I keep going?") — that is exactly the pause to avoid — and never to re-confirm a decision already made or to ask what a query would answer.


Worked example — "why is my plugin slow to open?" (the flagship)

This is the safest, most relatable case: bounded, one-shot, main-thread, deterministic — no real-time hazard, works regardless of the DSP story.

pulp trace start --categories render,gpu,text,js,layout
# ... open the editor ...
pulp trace stop
# Investigate the printed .pftrace using the offline query loop below.

A good answer reads like this:

Root cause: first editor-open spends ~2.4 s, and ~1.9 s of it is one-time GPU/text setup on the UI thread before the first frame — none of it cached between opens.

Breakdown: Dawn device + Graphite context init ~620 ms, Skia font-atlas build for the UI typeface ~540 ms, QuickJS eval of the bundled UI script ~410 ms, first Yoga layout + TextShaper::prepare of every label ~330 ms.

Chain of evidence: (1) a gpu/render span pair brackets the ~620 ms device init on the main thread. (2) text spans show the font-atlas build (~540 ms) preceding the per-label TextShaper::prepare. (3) the js span for script eval is ~410 ms, single-shot. (4) a second open repeats the same spans identically — nothing is reused.

Fix: warm the Dawn/Graphite context and font atlas once per process (not per editor open) and cache the compiled JS module. Re-opens should drop from ~2.4 s to well under 500 ms.

Editor-open recipe for a scripted / imported plug-in editor

Measure the open the way the host does it, and measure the full UI drawn, not just the view-creation call:

  1. Drive the real format path, off-screen, without stealing focus. For AU v2, load the .component in-process, ask kAudioUnitProperty_CocoaUI, call -uiViewForAudioUnit:withSize:, put the view in a borderless window ordered in with -orderFrontRegardless off every screen, Accessory activation policy, PULP_AUDIO_DEVICE=null. Timestamp: factory call (what the host blocks on), first drawable/present, content present (first present after the document mounted), idle (main thread answers a 2 ms heartbeat for 250 ms). Open 1 is cold; reopen N≥2 times for warm. Spectr's tools/editor_open_probe.mm is a working implementation.
  2. A/B with alternating rounds: best and median of ≥5 rounds per variant, alternating variants each round so host-load drift hits all of them alike. Report cold and warm separately.
  3. Capture with a traced SDK (PULP_TRACING=ON, kept under ~/.pulp/sdk-trace/), PULP_TRACE_PATH=… PULP_TRACE_RING_KB=524288 PULP_TRACE_SECONDS=60.
  4. Rank the phases of the Nth open (OFFSET picks the open):
SELECT ROUND((s.ts - r.ts)/1e6,1) t_ms, s.depth - r.depth d, s.category, s.name,
       ROUND(s.dur/1e6,1) ms
FROM slice s JOIN (SELECT * FROM slice WHERE name = 'scripted_ui_document_load'
                   ORDER BY ts LIMIT 1 OFFSET 1) r
  ON s.track_id = r.track_id AND s.ts >= r.ts AND s.ts <= r.ts + r.dur
WHERE s.dur > 2e6 AND s.depth <= r.depth + 3 ORDER BY s.ts;

Spans Pulp emits on this path: scripted_ui_document_load, scripted_ui_probe_realm (should be ABSENT on a first load — if present the document is evaluated twice), scripted_ui_live_realm, script_compile vs script_bytecode_read (a warm open should read, not compile), script_execute, runtime_import_parse / runtime_import_fonts / runtime_import_payload_eval / runtime_import_inline_eval, frame_callback_pump → raf_flush (the post-mount settle). runtime_import_verify (a child of runtime_import_parse) is the full decode + SHA-256 pass of a materialized document: it should appear on the first open only. On a reopen runtime_import_parse is a lookup well under 1 ms; a runtime_import_verify there means the document's bytes differ between opens (or PULP_RUNTIME_IMPORT_CACHE=0). Count, don't time:

SELECT name, COUNT(*) n, ROUND(SUM(dur)/1e6,1) ms FROM slice
WHERE name IN ('runtime_import_verify','script_compile','script_bytecode_read')
GROUP BY name;   -- 3 opens of one document: runtime_import_verify n = 1

Fingerprints and what they mean:

  • A slow COLD factory (view-creation call) whose span has no children while warm is fast: first-time app-side work in the plug-in's create_view(), not Pulp. Check the gate first — zero script_compile / script_execute / js_native / scripted_ui_* slices inside the factory span, against a non-zero count elsewhere in the same trace — then bracket the editor-create steps with temporary spans. A once-per-process sweep of the temp directory is a known case: the per-user $TMPDIR holds tens of thousands of entries on a developer Mac, and a remove_all of stranded packages adds to it. Running the same binary with TMPDIR= an empty directory is a one-variable A/B for it. Move such work onto a thread the plug-in image joins on unload, or after the view is returned.

  • Two copies of the whole load per open (compile → import → settle, twice): a probe realm evaluated the document. Fixed in Pulp for first loads; if you see it, the session is being reloaded, not loaded.

  • A long raf_flush / frame_callback_pump with trivial bridge calls under it: React work inside frame callbacks. Count commits, not natives: a LegacyRoot commits every setState made outside a batch. Pulp runs every rAF/timer callback through __pulpBatchUpdates__; a vendored bundle older than @pulp/react runtime revision 3 never installs it (pulp_check_vendored_react_runtime() says so). Promise-job and passive-effect commits are NOT batched by this — look for .then(setX) chains and mount effects that set state.

  • js_native spans that start a script-driven span corrupt nesting: __traceBegin__ inside a js_native scope closes on the native's end. For JS-level attribution of a bundled runtime, wrap functions with a timing accumulator (performance.now() deltas into a global table dumped with console.log from a setTimeout) instead of trace spans.

  • getLayoutBoxMetrics counts are a symptom, not the cost; measure the commit (see getLayoutRect coalescing in view-bridge).

Getting a traced SDK, and proving it is one

A plug-in or app only emits spans when it links a PULP_TRACING=ON SDK. Build one from the exact source you are measuring (a clean checkout at the release tag, or a committed branch — the build snapshots the commit, not the working tree):

git -C <pulp-checkout> checkout --detach v0.907.0     # or your committed branch
cd <pulp-checkout> && PULP_BUILD_JOBS=4 caffeinate -u -d -i \
  pulp sdk install --local --profile trace --print-path
# -> ~/.pulp/sdk-dev/trace-v1/darwin-arm64/<source-sha>/<fingerprint>

On a busy host the governor refuses a full-width request ("could not acquire build capacity"); PULP_BUILD_JOBS=4 asks for a share it can grant. Then prove it, by symbol count, not by the profile name: nm <prefix>/lib/libpulp-perfetto.a | grep -c perfetto (thousands) and nm <prefix>/lib/libpulp-view-script.a | grep -c perfetto (non-zero; a release SDK prints 0). Point the plug-in at it with -DPulp_DIR=<prefix>/lib/cmake/Pulp. A plug-in whose sources name a category the SDK does not declare (PULP_TRACE_SCOPE_NAMED("audio", …) — the list is the perfetto::Category(...) block in pulp/runtime/trace.hpp) compiles in a release build and fails only here, with kCatIndex_ADD_TO_PERFETTO_DEFINE_CATEGORIES_IF_FAILS_<line>; lint category literals against that header in the plug-in's own tests.

Editor open out of process (AUHostingService), with spans

Logic runs AU v2 plug-ins in AUHostingService, which inherits no environment. With a traced build, write ~/.config/pulp/trace-autostart:

PULP_TRACE_PATH=/Users/<me>/traces/oop/
PULP_TRACE_SECONDS=11
PULP_TRACE_RING_KB=262144

then drive opens with tools/editor-open/editor_open_oop_probe.sh (fresh instance per open; --gui-session over ssh) and read /Users/<me>/traces/oop/AUHostingServiceXPC_arrow-<pid>.pftrace. Keep the flush (PULP_TRACE_SECONDS) inside the probe's run: the service exits with its last client. Delete the file afterwards — every traced Pulp process records while it exists. The probe's "view controller N ms" is the host's placeholder time; the trace says what filled it. Order of work inside it: spectr_editor_create-style app spans (create_view), then editor_first_frame → scripted_ui_document_load (→ script_compile | script_bytecode_read, script_execute, runtime_import_verify, the app's bind span) → plugin_editor_frame (gpu_submit, first_frame_gpu_wait). A cold gpu_submit or first_frame_gpu_wait of hundreds of ms on the first open of a newly built/installed binary is Metal compiling shaders for that binary; re-run in a new process before attributing it to the editor. Counted, not timed, the prewarm proof is script_compile with no script_precompile before it (not prewarmed) versus script_precompile on another thread and only script_bytecode_read inside the factory (prewarmed):

SELECT t.name AS thread, s.name, COUNT(*) n, ROUND(SUM(s.dur)/1e6,1) ms
FROM slice s JOIN thread_track tt ON s.track_id = tt.id JOIN thread t USING (utid)
WHERE s.name IN ('script_precompile','script_compile','script_bytecode_read',
                 'runtime_import_verify','scripted_ui_prewarm')
GROUP BY 1, 2 ORDER BY 1, 2;

First-frame colour recipe: what did the editor show before its UI?

A content-first editor (the default for every hosted format; see view-bridge, "Editor open") mounts its document inside the host's view-creation call and presents it as frame 0. A view-first editor (PULP_EDITOR_OPEN=view-first, or an SDK that predates content-first) presents frames before its document mounts. Prove which happened from the trace, and what those frames showed from the pixels — the trace alone cannot tell a correct background from a framework default.

  1. Trace: the macOS plug-in GPU host emits one plugin_editor_frame (render) span per presented frame with args frame (restarts at 0 for each host, so every open begins at a frame-0 span), background_rgb (the colour painted under the tree), root_children, and width/height (the logical size the frame was laid out at). A content-first open wraps the mount, its settle rounds and frame 0 in one editor_first_frame (render) span inside the view-creation call; its scripted_ui_document_load therefore starts BEFORE frame 0. Pair each document load with the nearest frame 0, not the next one:
with frames as (
  select ts, dur, extract_arg(arg_set_id, 'debug.frame') as frame,
         extract_arg(arg_set_id, 'debug.root_children') as kids,
         extract_arg(arg_set_id, 'debug.width') as w,
         extract_arg(arg_set_id, 'debug.height') as h
  from slice where name = 'plugin_editor_frame'
), opens as (
  select ts as open_ts, row_number() over (order by ts) as n,
         lead(ts) over (order by ts) as next_open
  from frames where frame = 0
), dl as (
  select ts as doc_start, ts + dur as doc_end, dur as doc_dur
  from slice where name = 'scripted_ui_document_load'
), docs as (
  select * from (select dl.*, o.n, row_number() over (partition by dl.doc_start
                 order by abs(o.open_ts - dl.doc_start)) as rk
                 from dl cross join opens o) where rk = 1
)
select o.n as open,
       (select count(*) from frames f where f.ts >= o.open_ts and f.ts < d.doc_end
          and (o.next_open is null or f.ts < o.next_open)) as frames_before_document,
       round((d.doc_start - o.open_ts) / 1e6, 1) as doc_start_rel_frame0_ms,
       round(d.doc_dur / 1e6, 1) as doc_load_ms,
       (select printf('%dx%d kids=%d', f.w, f.h, f.kids) from frames f
          where f.ts >= o.open_ts order by f.ts limit 1) as frame0
from opens o join docs d using (n) order by o.n;

Content-first reads frames_before_document = 0, a negative doc_start_rel_frame0_ms and a frame 0 that already holds the document's children. View-first reads 1+ frames before the document with frame 0 at the chrome-only child count (Spectr: kids=1, its resize grip, vs 5 once mounted). In view-first, background_rgb reading #1E1E2E (kEditorHostClearRgb) means the plug-in declared no editor_background() and its root theme is the default dark one; every frame in frames_before_document showed that colour. Do not use root_children > 0 alone as "document mounted": chrome a processor adds in create_view() counts too.

Break editor_first_frame down with its children (depth <= +3, > 2 ms): scripted_ui_document_load → scripted_ui_live_realm → script_compile (cold) or script_bytecode_read (warm) + script_execute, the plug-in's own bind span, then frame 0's plugin_editor_frame with gpu_submit and first_frame_gpu_wait (the GPU finishing frame 0 before it is presented with its transaction). Measured, Spectr AU v2 in process: warm ~140–210 ms (document ~125, frame 0 ~15–35); cold ~390–640 ms, where a cold gpu_submit swung from 23 to 325 ms between runs on a loaded host — read cold numbers as a distribution, alternate variants, and repeat.

  1. Pixels: read back every presented drawable in the host process (swizzle -[CAMetalDrawable present] / -[MTLCommandBuffer presentDrawable:], blit the texture into a shared MTLBuffer) — never a screenshot, which needs screen-recording permission and sees the compositor, not the frame. Also read the view's layer.backgroundColor before the window is ordered in: that colour is on screen until the first frame and no frame read-back sees it. Stamp each present with clock_gettime_nsec_np(CLOCK_UPTIME_RAW) — Perfetto's clock on macOS — and a present lands within ~1 ms of its plugin_editor_frame span's end, so frames join spans without guessing.
  2. Classify every frame, not a sample: the editor's background (plus chrome already in its final place), the settled UI, the framework default, or other. Report counts per class and how long an off-brand frame stayed on screen (until the next present). Run it cold and warm, at the preferred and minimum host sizes, and on a build without the declaration as the negative control. A traced SDK draws a TRACING badge in the editor's corner: expect those pixels in the classes; gate on an untraced build. Geometry is part of the gate. A frame laid out at a size other than the host's — the preferred-size layout cropped into a minimum-size window, or a new-size drawable scaled into the layer's old bounds — fails like an off-brand colour: every presented frame's layout size must equal the view's bounds and the host's size. It shows only when the host size differs from the preferred size, so the minimum-size run is the one that can catch it, and out of process (AUHostingService) is where it happened. In a trace, the first plugin_editor_frame after scripted_ui_document_load ends must carry the host's width/height, not the preferred size, when the host resized during the mount. In the host-window read-back, check the image's own size first: a capture whose dimensions differ from the requested host size is the host's window still converging (seen at the preferred size with and without the plug-in's fix), not the plug-in's frame; report it as its own class. The in-tree repro is test_plugin_view_host_first_frame_macos.mm ("a host resize queued during the mount lands before the next frame"); the rules it pins are in the view-bridge skill.
  3. Count the stages a user sees from the host window, not from the presents: read back every distinct image of the host's own window from the moment it asks for the editor (magenta backdrop), and count the images before the settled UI that are neither that backdrop nor already the UI. Run once with AppKit's window animation and once without (NSWindowAnimationBehaviorNone); with it, smaller-than-final images are expected, so classify them by block means against the settled image (the UI being zoomed in is one stage; the empty background being zoomed in is the "small, then empty, then UI" report). Target: 0 non-UI images per open.
  4. Keep the display awake for any off-screen probe: caffeinate -u -d -i <probe>. A sleeping display stops the display link, so every variant presents zero frames and the run reads as a broken build.

The other canonical case is the offline DSP reveal — "CPU pinned but the meter looks calm — which node?" — run against a deterministic offline_process() render (examples/trace-demo) so the answer reproduces exactly. See hints_dsp.md and docs/guides/tracing.md use case 3.


Agent contract

  1. State the question as a measurable target before querying.
  2. Verify the capture is non-empty and has the categories you need before analyzing. An empty trace is a capture bug, not a "fast" program.
  3. Distinguish wall time from CPU time before blaming a long span (step 3).
  4. Follow blockers across threads; the slow thread is often not the guilty one.
  5. Verify globally — report the ranked budget, not the first bottleneck.
  6. Every claim cites a query result. No evidence, no finding.
  7. Investigate autonomously; AskUserQuestion only at a genuine priority fork, never for something a query answers or a progress checkpoint.
  8. L1 returns prose (root cause + evidence + fix); never dump SQL at a novice.

Files this skill covers

  • .agents/skills/trace-analysis/references/hints_dsp.md
  • .agents/skills/trace-analysis/references/hints_frame.md
  • .agents/skills/trace-analysis/references/hints_js.md
  • .agents/skills/trace-analysis/references/hints_gpu.md
  • .agents/skills/trace-analysis/references/hints_gpu_audio.md
  • .agents/skills/trace-analysis/references/hints_crossplatform.md
  • .agents/skills/trace-analysis/references/ui_jank_playbook.md — live-editor jank workflow
  • .agents/skills/trace-analysis/references/ui_jank.sql — its PerfettoSQL definitions
  • .agents/skills/trace-sql/SKILL.md — the SQL substrate + trace-stdlib
  • core/runtime/include/pulp/runtime/trace.hpp — macro surface + category taxonomy
  • docs/guides/tracing.md — the guide, tiers (L0/L1/L2), worked use cases, gotchas

Tracing a plug-in on Windows

Windows tracing was unusable until 2026-07-25 — four independent blockers, each fatal on its own. If a Windows capture comes back empty, check these first.

  1. PULP_TRACING=ON did not compile under MSVC. trace.cpp's ship-guard sentinel used __attribute__((used, visibility("default"))); MSVC errors C4430/C2065/C3861 and pulp-runtime fails, taking 22 dependent targets with it. Now __declspec(dllexport) on MSVC.
  2. Perfetto was excluded from the installed SDK. pulp-runtime linked tracing through $<BUILD_INTERFACE:pulp-tracing>, correct when tracing is OFF and wrong when ON — every library carries Perfetto symbols but the export named none, so a plug-in linking a traced SDK failed with unresolved perfetto:: symbols. pulp-perfetto/pulp-tracing/perfetto.h now install into PulpTargets when tracing is ON.
  3. No code path ever started a session. Tracing::start() had zero callers and a plug-in has no main(). Tracing::attach() now autostarts from $PULP_TRACE_PATH, and the VST3 adapter attaches/detaches over its lifetime.
  4. The Windows plug-in host had no spans. Only window_host_mac.mm had render/frame + canvas/paint.

Capture recipe

cmake -S <pulp> -B <build> -DPULP_TRACING=ON       # then build + cmake --install
# in the HOST process environment (not the build shell):
PULP_TRACE_PATH=C:\path\out.pftrace
PULP_TRACE_SECONDS=45      # timed flush; see below

PULP_TRACE_SECONDS matters. Perfetto's duration_ms only caps the buffer — the .pftrace is written by stop(), which otherwise means unloading the plug-in you are profiling. The timed flush makes a capture self-completing.

For an SDK-built plug-in, prove that the installed package was configured with PULP_TRACING=ON; enabling the option only in the plug-in's consumer build cannot reconstruct omitted Perfetto targets or headers. A valid traced SDK exports the tracing support targets transitively, and the host process—not the build shell—must receive PULP_TRACE_PATH and PULP_TRACE_SECONDS. If the plug-in loads but produces no file, distinguish an untraced installed SDK from a session that merely has not flushed before changing instrumentation.

The auto-flush timer is owned, joined, and generation-tagged

PULP_TRACE_SECONDS used to arm a DETACHED std::thread that slept and then called back into process-global tracing state. Two consequences you may still see in older builds:

  • A capture that truncates early. Close and reopen the editor inside the window and the FIRST session's timer stopped the SECOND session. Timeouts now carry the session generation they were armed for and refuse to act on any other, so a re-opened editor gets its full window.
  • A crash on plug-in unload. Nothing joined the sleeping thread, so FreeLibrary / dlclose could pull the module out from under it. The final Tracing::detach() now cancels and JOINS the timer before it flushes.

Practical consequence for capture: the last detach is a synchronous flush + join. If you are scripting a capture, let the host finish unloading the plug-in rather than killing the process — a SIGKILL still loses the trace, but a clean unload no longer races the timer.

Adapters attach via RAII (runtime::ScopedTracingAttachment), and tracing is now wired into VST3, CLAP, AU v2, AU v3, AAX, and Standalone — it used to be VST3-only, so a Perfetto capture of any other format recorded nothing while the API claimed to be process-global. If a capture is empty, check the format is one of those before suspecting the environment.

Always instrument the blocking call

A frame span whose children sum to ~2 ms while the frame itself takes 45 ms means the cost is in an uninstrumented call inside it. On Windows that was the swapchain acquire (gpu_acquire, added 2026-07-25): with a Fifo present mode GetCurrentTexture() blocks until the next refresh. Before that span existed the time had nowhere to be attributed and the trace looked healthy.

Reading a long gpu_acquire on macOS: CPU bunching or GPU-bound?

On Metal the acquire is [CAMetalLayer nextDrawable], which blocks until one of the layer's drawables (three by default; Dawn and Pulp set neither maximumDrawableCount nor allowsNextDrawableTimeout) comes back. A wait of one or two refresh intervals means every drawable was held, and in the standalone GPU window the span's args say by whom (a plug-in editor's gpu_acquire carries no args yet):

argmeaning
frames_in_flightframes submitted to the GPU and not yet finished (as of the last completion pump; may read one high)
gpu_render_mslast sampled GPU render time — 0 unless GPU timing is on (PULP_GPU_TIMING=1 for a standalone window)
late_msacquire start minus the display-link target presentation time; negative = rendering ahead
refresh_period_msthe display's refresh interval
vsync_drivenfalse for resize / capture / first-show frames rendered outside the link

frames_in_flight ≥ 2 with gpu_render_ms near or above refresh_period_ms is a GPU-bound frame: cut GPU work. frames_in_flight ≤ 1 with the previous frame having presented inside the same refresh interval (negative or small late_ms on back-to-back frames) is CPU bunching: two presents landed in one interval and filled the queue.

Known issue — "nonblocking" is not non-blocking on macOS. Dawn's Metal backend treats Mailbox exactly like Fifo (it can only toggle displaySyncEnabled, which only Immediate turns off). The plug-in editor's PresentPolicy::nonblocking asks for Mailbox first, so on macOS it is still paced to vsync and can still block in acquire. Do not read a macOS editor's acquire wait as proof the policy is broken elsewhere.

Measuring without touching the user's audio. A standalone launched with PULP_AUDIO_DEVICE=null renders its audio graph on a real-time paced thread with no output device (PULP_TEST_SIGNAL still feeds the input), so a live-window trace can run with meters and analyzers publishing at their normal rate and nothing reaching the speakers.

Driving a Windows GUI capture with nobody watching

Screenshot/input automation needs an Active session; a disconnected RDP session captures blank frames at the default 800x600. Move the session to the console so it stays renderable with no client attached:

tscon <session-id> /dest:console     # session stays Active, RDP client detaches

Pair with auto-logon so a session exists after reboot, and run long builds under Task Scheduler (S4U) — an SSH drop otherwise kills cmake/MSBuild mid-build.

A capture with NO render spans — read this before investigating

A trace containing layout_children and wm_mousemove but no frame, paint, gpu_acquire, gpu_submit or gpu_present is the single most misleading result this harness produces. It looks like a broken capture or a frame-time regression and is usually neither. Work the ladder in order; each step is seconds, and step 1 explains most cases.

1. Which host did the plug-in actually get?

The five render spans live on the GPU paint path. A plug-in that does not ask for a GPU editor never enters it, paints correctly on CPU raster, and emits none of them — by design, on every platform.

The adapter logs its decision at editor attach:

[plugin-gpu-host] adapter mode=autoui use_gpu=false wants_gpu=false
                  scripted=false requires_gpu_host=false …
VST3 editor: attached (536x230, mode=autoui, gpu=false)

decide_gpu_host() (core/format/include/pulp/format/gpu_host_select.hpp, no platform guards — this is cross-platform) computes:

d.wants_gpu = scripted || view_wants;      // view_wants = requires_gpu_host()
d.use_gpu   = d.wants_gpu && !env_off;

So mode=autoui with wants_gpu=false means the editor is neither scripted nor declares requires_gpu_host(). That is a complete explanation for zero render spans. Do not go looking for a fallback, a broken adapter or a regression: a CPU-raster editor is not a degraded GPU editor, and its frame times are not comparable to a GPU capture's. mode= is the first thing to read.

On Windows the CPU branch is explicit — handle_wm_paint() calls render_frame() only if (gpu_surface_ && skia_surface_), else raster_render_rgba(), which carries no instrumentation and is labelled in source as "the VM proof path".

2. Which host + platform emits which spans?

Coverage is not uniform, and only the Windows plug-in editor has the full set. Verified by span-site inspection:

hostframe / paintgpu_acquire
Windows plug-in editor (plugin_view_host_win.cpp)yesyes (shared PluginFrameRenderer)
Linux plug-in editor (plugin_view_host_linux.cpp)noyes (shared renderer)
macOS plug-in editor (plugin_view_host_mac.mm)nono — it has its own render_frame() and no PULP_TRACE sites
macOS standalone app (window_host_mac.mm)yesyes (own span site)

gpu_submit comes from core/render/src/skia_surface*.cpp and gpu_present from gpu_surface_dawn.cpp, i.e. the render layer rather than the host, so they can appear where the host-level spans do not.

That ownership is exclusive: a host must NOT bracket skia_surface_->end_frame() / gpu_surface_->end_frame() with a span of its own. The surface already opens one on every path, so a host-level wrapper emits the stage twice per frame — and only on the paths that reach it, since a bail-out that presents without submitting skips the wrapper entirely. The per-frame count then differs between paths, defeating any contract that counts one span per stage per frame. gpu_acquire is the one stage a host does own: Dawn's begin_frame() opens no span.

The practical consequence: do not read a missing frame span on a macOS plug-in editor as a regression — that host has never emitted one. Instrument the host before measuring it, or measure the standalone app instead.

3. Read the plug-in's log before theorising

On Windows Pulp's log sink is OutputDebugStringA, which nothing captures by default — so [plugin-gpu-host], GPU-init failures and the CPU-fallback diagnostic are all invisible unless you attach a listener FIRST.

There is no need for Sysinternals: a DBWIN listener is ~40 lines. Create the DBWIN_BUFFER file mapping plus the DBWIN_BUFFER_READY / DBWIN_DATA_READY events, then loop SetEvent(ready) → WaitForSingleObject(data) → read the pid from the first 4 bytes and the message after it. Start it before the host process and it captures everything from plug-in load onward. This is what turns "no spans, cause unknown" into one line of fact.

On macOS the same information comes from log stream — see the ios skill for a working predicate that includes [plugin-gpu-host].

4. Only then suspect the capture

If mode=scripted/custom with use_gpu=true and the spans are still missing, the earlier Windows blockers and the flush-lifetime notes above apply.

Related but different: the section below concerns gpu_render_time, an opt-in timing COUNTER. Missing gpu_render_time and missing render SPANS have unrelated causes; do not treat one as evidence about the other.

GPU render time is now OPT-IN (WAH-13)

SkiaSurface::gpu_render_timing_available() reporting false is no longer evidence of an adapter that lacks timestamp-query. Timestamps are requested only when the host asks, via PluginViewHost::Options::enable_gpu_timing (default OFF), rather than whenever the adapter advertises the feature.

That default is deliberate and worth understanding before you "fix" it: Dawn gates writeTimestamp behind the allow_unsafe_apis toggle on every backend, so requesting the feature forces that toggle on — and it applies to the DEVICE, not to the diagnostic. Ordinary rendering was silently running with relaxed validation on every machine whose adapter happened to offer timestamps.

If you need per-recording GPU time in a capture, enable it explicitly on the host's Options. If a trace shows no gpu_render_time, check that flag before suspecting the adapter.

GPU errors arrive as gpu.diagnostic instant events

A Dawn uncaptured-error / device-lost callback and a Skia log record no longer live only in runtime::log_error output. core/render/src/gpu_diagnostics.cpp forwards them into the timeline as zero-duration instant events named gpu.diagnostic on category gpu, with the severity and the message text riding along as debug annotations. They keep their log_error calls, so a log and a trace should agree — a diagnostic in one and not the other means the bridge was off, not that the event did not happen.

The Skia half is gated: it installs an SkLogHandler only when tracing is compiled in or PULP_GPU_LOG_BRIDGE is set, and only when nothing else already owns Skia's process-global handler. A capture with Dawn diagnostics but no Skia ones is therefore an ordinary outcome (a host already held the slot), not a dropped event. skia_log_bridge_state() reports which it was.

The log copy of a Skia record is capped (kSkiaLogTextLimit, 50 per process) while the trace event is not, so past that cap a trace legitimately holds Skia records the log does not — the GpuDiagnostics: … skia_text_suppressed=N line a headless screenshot run prints says how many. Compare skia_records on that line against your gpu.diagnostic count before calling the two out of sync.

DPR experiment traces reuse A2T

A4 DPR trials do not introduce a second profiler or a new ad-hoc SQL report. They must answer the named A2T Perfetto questions and carry the protected A3 policy/campaign identity required by docs/contracts/gpu-dpr-experiment-v2.schema.json. V1 receipts are historical nonterminal input and count as zero v2 cells. Validate required categories from the DPR manifest before interpreting a run. Missing categories, unavailable real GPU timing where the question requires it, or a trace that cannot be bound to the trial artifact leaves the cell incomplete; it is not evidence that the cost was zero.

tools/scripts/gpu_dpr_runner.py validates each cell's raw trace artifact, required categories, and named gpu-startup, gpu-health, and gpu-probe receipts before ingestion. It does not run substitute SQL or turn missing GPU timing into zero; missing or rejected trace evidence remains resumable. The v2 runner first snapshots the exact A2T-authorized analyzer and each producer trace into its owned run root. It hashes held regular files, rejects symlink/outside/substituted paths, and reruns that analyzer during ingestion, finalization, and complete-result validation. JSON text that merely names trace categories is not a Perfetto trace; browser cells must retain their real DevTools trace form. The state HMAC is integrity-only and cannot replace the nonce-bound trace receipt or protected-main dependency blobs. Each named answer must expose one category_scope matching the trial evidence ID and stable Perfetto process instance across all three questions. Categories from another PID/UPID or nonce do not satisfy the manifest. Treat the closed probe failure diagnostics cpu_oracle_mismatch and magnitude_dispatch_failed as causal failures even if adapter health was reported healthy.

Each v2 mode emits 30 aligned trials of at least 240 frames in the same process as the correlated trace, plus five independently reset warm-ups. Its 20 first-frame values are different: each comes from a fresh child process and has a typed ledger row binding attempt nonce/number, unique PID, producer/content/build digests, exact adapter identity, and the sample. The adapter and runner revalidate that ledger; a reused PID or mixed identity is incomplete evidence, even when the aggregate sample count is 20. Repeat the complete 84-cell matrix on the same frozen machine/provider/build; an isolated attractive trace cannot replace the same-unit repeat formulas.

Trace timing does not bypass instrument validity. The raw metric must say whether it is measured, derived, or unavailable, and GPU timer evidence must include an empirical resolution plus five baseline and five eight-times-work calibration trials whose medians are distinguishable. Adaptive traces must bind the actual measured samples and scale-before/scale-after transitions; a mode label or requested scale alone is not observed adaptive behavior.

Correlating GPU-health startup snapshots

The Pulp-owned product host creates distinct bounded GPU and trace evidence IDs and emits both on its first-visible lifecycle and health-transition spans. A role adapter must preserve those exact values in the health response and trace; it must not synthesize a replacement when either is absent. Keep causal_attribution=unverified and the startup verdict unverified until an A2T artifact supplies the required categories and is bound to the same product instance, build, frame identity, and trial. Dropped events, truncation, missing categories, timeout, or instance loss are fail-closed evidence, not zero-cost measurements. Perfetto production and provider mutation remain off the audio thread.

For the agent slash surface, use one named question: /trace gpu-startup --trace FILE, /trace gpu-health --trace FILE, or /trace gpu-probe --trace FILE. Preserve the corresponding pulp trace ... --json output for the A2T receipt. A human Perfetto UI inspection is useful additional evidence but does not replace the digest-bound machine result or its invalid-trace negative.

For an A3 external role, the executable selected by PULP_A3_CAMPAIGN_PRODUCER must emit the campaign trace with the exact health IDs. The checked-in external adapter snapshots that producer but does not run substitute SQL or create an evidence ID. Its role producer pins the exact checked-in source-bound analyzer wrapper, prepares it with fresh config-free Cargo state and retained toolchain/source/output digests, proves its invalid-trace negative, and derives the typed campaign analysis only when the named replay selects the health result's exact evidence ID and the producer-challenged trace-host PID. The structural replay is normally unverified/exit 2 because it has no A3 budget; keep it separate from the health document's budget verdict. A failing analyzer result is terminal, and the producer may not overwrite either verdict. Final same-instance A2T replay and human Perfetto UI correlation are added only from the selected causal campaign after all role runs complete. The canonical v2 matrix has exactly seven entry roles: pulp-standalone, forge-modular-standalone, forge-modular-auv2-logic, forge-modular-vst3-reaper, forge-modular-clap-reaper, headless-reference, and constrained-adapter. Each entry point pins its lifecycle driver and rejects analyzer evidence IDs that differ from the exact health instance. The driver still owns capture; the entry point cannot manufacture a Perfetto file or native-present event. Do not collect until the protected product policy binds all role thresholds and the constrained adapter, configuration, support matrix, and authentic A1 evidence; the current canonical receipt is truthfully blocked on those two inputs.

The DPR native adapter is snapshot-executed, so it must stay stdlib-only at module scope. gpu_dpr_runner.run_cells copies the adapter alone to run_dir/tooling/adapters/<key>/<nonce> and executes that lone copy, so a module-scope sibling import such as import gpu_dpr_evidence passes every in-tree test and still breaks the product path.

The product health-transition spans are real runtime producers even though the macros compile out with PULP_TRACING=OFF. Terminal A3 separately requires the four-state pre-change/compile-out/compiled-in-idle/active control documented in docs/validation/gpu-first-visible-a3-acceptance.md: 5 warmups, 30 measured, 20 fresh-process observations per state, zero xruns/audio-thread trace events, and derived median/p95 ceilings. The offline A2T no-producer disposition does not waive this product control.

Collect every state with gpu_first_visible_a3_trace_producer_overhead.py collect-state, never by invoking the product driver directly. The collector owns 55 live challenge/ack handshakes per state and binds exact executable/start identity. Active replay must prove both immutable packages: health-first-visible at 8175bd… and the complete b4ba exact 20-signature state/render/js inventory. Acquire/submit/present are mandatory per sample; the other 17 signatures are counted and an unobserved signature stays explicitly not-covered, never zero-cost. Use the exact v57.2 trace_processor_shell; production Chrome JSON is forbidden. The request also pins a candidate-relative state_build_driver. Before timing, the collector exports the exact row source and default-deny rebuilds it without the measured binary, ambient build output, or network; rebuilt bytes and the compile-in sentinel must match the requested state. Preserve the source archive, closed build request/receipt, product, logs, and digest/version-bound toolchain snapshots. A direct binary or build-driver assertion cannot pass.

The evidence requirement is question-scoped

gpu-startup answers for a capture taken with no instrumentation at all: when no span anywhere in the trace carries an evidence id, it admits a single untagged cohort and reports the setup work the capture plainly contains, with a NULL evidence id. The gate is the whole trace, never the startup candidates alone; a capture whose gpu_probe* spans are tagged is instrumented, and is not admitted. gpu-health and gpu-probe do not relax — an untagged capture stays unavailable there, because each of those answers is a correlation claim, and a correlation with nothing to correlate is not a weaker answer but a different one.

Read unavailable_reason before you conclude the capture is uninstrumented. Those two questions separate the two ways their answer can be empty: invalid-evidence-correlation means the capture carries the question's spans and they could not be correlated (untagged, mixed, or malformed evidence), while missing-question-category means the capture carries none of that work at all. The first is a recapture-with-correct-evidence problem; the second is a recapture-with-the-right-categories problem, and treating one as the other sends you to re-record a trace that was already recording the right thing. Only gpu-startup still answers both with missing-question-category, since it admits an untagged cohort and an empty startup answer is not a refusal.

Two things follow. A gpu-startup breakdown is not evidence that the capture is tagged, so it does not predict that gpu-health or gpu-probe will answer at all. And the relaxation is all-or-nothing about instrumentation: gpu-startup admits the untagged cohort only when no span anywhere in the trace carries a debug.gpu_evidence_id. One evidence id — on a startup span, or only on a gpu_probe* / gpu_readback* / gpu_health_transition span that is not a startup candidate at all — drops the untagged cohort and puts gpu-startup back on the exact shared-evidence-id requirement.

Absence of evidence is not on its own enough to answer, because an untagged cohort has no id to separate one lifecycle from the next. gpu-startup also fails closed on an untagged capture holding more than one frame-zero anchor, or spanning more than one process, and reports unavailable / missing-question-category with no contributors — the same observable answer the tagged path gives a capture with two first-visible lifecycles. So an untagged gpu-startup answer is bounded by at most one frame-zero anchor in one process. That bound is weaker than the tagged path's, which separates lifecycles by id: a second lifecycle carrying no frame-zero anchor of its own cannot be told apart from concurrent setup work, so it is reported with the rest rather than refused.

Correlate a catalog recipe with Perfetto

Begin with pulp gpu recipes list --symptom <exact-token> --json, run the selected baseline twice, then its seeded negative control. Carry the emitted gpu_evidence_id into the closed gpu-probe trace question so a scheduling or frame contributor can be tied to the exact oracle evidence rather than to a filename guess. Exit 2 remains unavailable/unverified and cannot be interpreted as zero cost. The separate exact-instance dev.pulp.gpu/health.read@1 control operation is a cheap live snapshot; it neither starts a trace nor proves an offline catalog recipe callable.

Perfetto localizes a stage; it does not automatically prove a platform event race. A useful trace may show inexpensive paint work (about 1 ms) alongside resize spans clustered near a 60 Hz frame (about 15.7 ms p50 and 18.7 ms p99), which focuses the next test on acquire, present, compositor, or callback order. Then use a deterministic AppKit/GPU event-order harness and a planted old- behavior negative control to prove whether a redundant same-size resize or retained-cover lifetime caused the symptom. Finish with real product proof, such as a 60 fps recording plus interaction/feel validation. Report those as three distinct links: trace localization, platform-race proof, and product acceptance. Never describe Perfetto alone as having found the root cause.

For A3 v2, do not validate a submitted analysis sidecar in isolation. Rehash the exact .pftrace, verify prepared-analyzer provenance, run that analyzer, and require its evidence/process scope and capture completeness to agree with the digest-bound campaign, category set, and sidecar bindings.

A script event handler is not opaque any more — but you must ask for it

dom_event_evaluate wraps the whole of a script's event handler, and for a long time nothing inside it emitted a span. Re-attributing a band-drag capture by self time (a slice's total minus the sum of its direct children) put 95.8% of that block outside every span the tree emitted — 2119.8 ms of self time under 2213.6 ms of total, across 127 pointer events. That is not a slow handler you can locate; it is a handler you cannot see into at all.

Two instruments now open it, and they answer different questions:

  • js_native spans — every JS→C++ native is registered through the one register_bridge_function template, so each call is wrapped in a js_native span (category js) whose debug.fn arg names the bridge function. The span name is the constant js_native; group by EXTRACT_ARG(arg_set_id, 'debug.fn'). A GLOB 'js_native:*' filter — the older per-function naming — matches nothing and returns a silent zero. This is what to reach for when the script is one this repo does not own (an imported design's runtime.js, a materialized React bundle): you get the native half of the handler attributed by name with zero edits to the script. It is compiled out when tracing is off.
  • pulpTrace.begin/end/scope(name, fn) (raw: __traceBegin__ / __traceEnd__) — a script naming its own spans. Only useful when you can edit the script, and only covers what you chose to wrap.

Read the two together. A handler whose self time collapses once js_native spans appear was spending its time in bridge calls; one whose self time stays high is spending it in the script's own interpreted work, and no native span will ever show you that — you need pulpTrace scopes in the script itself.

js_trace_force_closed_unbalanced_scope in a trace is a defect marker

pulpTrace pairing belongs to the script's control flow, so an early return or a throw between begin and end leaves a span open — which silently re-parents every later slice under a span that never closed, and quietly corrupts every attribution downstream of it. The dispatch boundary force-closes what a handler leaves open and says so three ways: the counter track js_trace_unbalanced_scopes, a slice named js_trace_force_closed_unbalanced_scope, and a line on stderr. If you see either in a capture, the trace's parentage before that point is suspect — fix the script's pairing (prefer pulpTrace.scope(), which is try/finally) and re-capture rather than reasoning about the numbers you have.

__traceStats__() returns {depth, forceClosed, unmatchedEnd, refused} from JS for the same reason, and those counters increment whether tracing is compiled in or not — so a test can assert balance on the default gate build where every Perfetto macro expands to nothing.

Before you read any number: prove the capture is not truncated

A Perfetto in-process capture fails in a way that looks exactly like success. The ring is fixed-size; when it wraps, the interned string table at the head of the sequence is overwritten, and every packet on that sequence after that point becomes unparseable. trace_processor loads the file without complaint and reports zero slices. What you have is a large .pftrace — a real, 100 MB file with a plausible timestamp — that contains nothing, and a query returning no rows against it reads identically to "that span never fired".

So a zero-row result is only a finding once you have shown the trace is intact. Run this first, on every capture, before quoting anything from it:

SELECT name, value FROM stats WHERE name IN (
  'traced_buf_write_wrap_count',
  'traced_buf_bytes_written',
  'traced_buf_bytes_overwritten',
  'traced_buf_buffer_size',
  'traced_buf_incremental_sequences_dropped',
  'packet_skipped_seq_needs_incremental_state_invalid')
  AND value != 0;

traced_buf_write_wrap_count > 0, any traced_buf_incremental_sequences_dropped (a whole sequence — often the main thread's — was discarded), or any packet_skipped_seq_needs_incremental_state_invalid means the capture is truncated — re-capture with a bigger ring, do not analyse it. Pair it with a positive control that must be non-zero for a capture of that workload (for a UI drag: SELECT COUNT(*) FROM slice WHERE name='dom_event_evaluate'), because a clean stats table on an empty trace only proves nothing overflowed.

The ring size the env-driven autostart uses is $PULP_TRACE_RING_KB (KB, default 80 MB, accepted range 1 MB–4 GB; a malformed value is refused on stderr rather than silently falling back). The 80 MB default is sized for a render/DSP capture. A script-heavy UI capture with js_native spans enabled will overrun it — one 6-second Spectr band drag wrote 100.4 MB into the 80 MB ring and produced a zero-slice file. Budget ≥ 256 MB (PULP_TRACE_RING_KB=262144) for that shape of capture.

Capturing jank that appears only during interaction

Frames that stall only while the user moves the mouse need a capture that isolates the interaction from everything else in the process:

  • Drive a real mouse and a deterministic animation together. Post kCGEventMouseMoved events at 60 Hz across the target while a fixed animation source (an LFO) runs, then compare a during-sweep window with a same animation, mouse still window in the same trace. The difference is the interaction's cost; anything present in both is not. A global CGEvent mover moves the real cursor and drives whatever window is under it, so use it only in an attended session; unattended, use PULP_TEST_POINTER_DRAG (next section).
  • Never capture while a scripted-scenario harness is stepping. Its per-step snapshots add 100–800 ms stalls that read exactly like app jank.
  • Size the ring and let it flush. Use PULP_TRACE_RING_KB=524288 for a script-heavy UI, and let PULP_TRACE_SECONDS plus the flush elapse before ending the process — an earlier kill writes no file.
  • Read the slices by side. dom_event_dispatch, dom_event_evaluate and __flushTimers__ name the JS-side cost; gpu_acquire the GPU wait; paint the drawing. Steady-state gpu_acquire of a few ms with ~2 ms paint and a dom_event_evaluate of tens of ms points at the script, not the renderer.

In a captured/materialized React import, a fat dom_event_evaluate on pointermove is almost always a React commit per move re-applying import metadata; the fix and its checklist are in the view-bridge skill ("Realtime scripted editors: the performance checklist").

Frame pacing of a live editor: recipe and environment traps

For an editor that animates live audio data while the user drags or zooms — the design-time rules it should already follow are the checklist in the view-bridge skill ("Realtime scripted editors: the performance checklist"). The full query workflow — health, per-phase frame gaps, worst-frame classification, content cadence — is references/ui_jank_playbook.md; this section is the short form.

Capture. Put everything in the launched standalone's environment:

PULP_TRACE_PATH=/tmp/run.pftrace \
PULP_TRACE_SECONDS=60 \
PULP_TRACE_RING_KB=1572864 \
PULP_TEST_SIGNAL=noise \
PULP_TEST_POINTER_DRAG='rect:0.20,0.50,0.80,0.50,120,4' \
  ./build-trace/.../YourPlugin.app/Contents/MacOS/YourPlugin
  • Size the ring for audio, not just script. A 60 s run with audio playing through a scripted editor writes on the order of 700 MB — nearly nine times the 80 MB default, and more than a 512 MB ring holds. A wrapped ring drops whole sequences, often the main thread's, and the file still opens; check stats (above) on every capture and shorten PULP_TRACE_SECONDS rather than accept a wrap.
  • Real audio through the plugin. PULP_TEST_SIGNAL=noise (or sine) feeds the standalone's input so meters, analyzer and signal-scaled effects run at their in-use cost. An idle editor measures nothing that matters.
  • Input inside the window. PULP_TEST_POINTER_DRAG injects the gesture into the window host's own mouse path on its frame schedule (rect:X0,Y0,X1,Y1,N[,R], normalized top-left coordinates; an unparseable value disables the drive). It cannot touch other applications' windows.

Compute frame gaps per phase. Take the interval between consecutive frame slices (render category) on the editor's thread and report median and p95 for each phase — idle with audio, mid-stroke, and the 2 s after each stroke's release — never one whole-run mean. The release window is where regressions hide: in one measured spectrum editor mid-stroke averaged 31 ms (p95 39 ms) while the 2 s after release averaged 56 ms (p95 184 ms). A commit at pointerup, a data backlog draining after the gesture, or a deferred relayout all land there.

Validity gates — discard any run that fails one:

  • the app was frontmost for the whole capture;
  • a window screenshot taken during the run shows the editor drawing;
  • the trace is non-empty, stats shows no wrap or dropped sequence, and it contains both editor paint slices and data-delivery activity (analyzer frames reaching the page) — the positive control for this workload.

Environment traps that fake a regression or hide one:

  • Locked screen or sleeping display → no display-link frames. The trace is well-formed and shows an idle editor.
  • A GUI app launched from a sandboxed agent shell may never attach to the window server. It runs and traces audio but paints nothing; the screenshot gate catches it.
  • CoreAudio itself can be degraded. One host accumulated 4,096 duplicate com.apple.AirPlayXPCHelper HAL plug-in objects, and every audio app — not just Pulp — spent roughly 19–75 s in HALSystem::InitializeDevices at launch. Diagnose by timing a kAudioHardwarePropertyDevices query and counting classes among kAudioObjectPropertyOwnedObjects of the system object; thousands of one class is the tell. The fix is sudo killall AirPlayXPCHelper then sudo killall -9 coreaudiod, or a reboot — a human's call, not an agent's.
  • Shared-host load skews timings more than most code changes do. Capture baseline and candidate back to back on the same host with the same signal and gesture; never compare against a number from another session.

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

同じリポジトリのスキル

概要と使いどころ

aax

無料

Optional AAX support for Pulp, including developer-supplied Avid SDK setup, CMake enablement, DigiShell/AAX Validator workflows, and local AAX builds on macOS or Windows.

日本語の概要は準備中です。原文の説明を表示しています。

Generous-Corp/pulp222026年10月10日 更新

Configure, implement, and test Pulp's optional desktop Ableton Link tempo-sync adapter while preserving the developer-supplied SDK, licensing, realtime, latency-compensation, and no-install boundaries.

日本語の概要は準備中です。原文の説明を表示しています。

Generous-Corp/pulp222026年10月10日 更新

Maintain Pulp's installed design-time agent capability manifest and public-surface ledger. Use when adding, removing, renaming, or materially changing public audio, MIDI, signal, timebase, or sequence APIs; registering a new algorithm for generators; changing capability support or deprecation state; or repairing agent-capabilities freshness, schema, fingerprint, tombstone, or installed-SDK tests.

日本語の概要は準備中です。原文の説明を表示しています。

Generous-Corp/pulp222026年10月10日 更新

android

無料

Android platform development for Pulp — NDK cross-compilation, Oboe audio, Dawn/Skia GPU rendering, JNI bridge, touch interaction, emulator workflows, and end-to-end smoke validation. Covers build, deploy, debug, and the gotchas discovered during bringup.

日本語の概要は準備中です。原文の説明を表示しています。

Generous-Corp/pulp222026年10月10日 更新

ara

無料

Optional ARA support for Pulp, including developer-supplied ARA SDK setup, CMake enablement, adapter companion APIs, validation, and ARA-aware plugin implementation guidance.

日本語の概要は準備中です。原文の説明を表示しています。

Generous-Corp/pulp222026年10月10日 更新

The measurement surface for ALL Pulp DSP and audio-pipeline work — read it BEFORE writing or gating DSP, not only when something already sounds wrong. Covers the C++ harness (signal generators, metrics, assertions, RenderScenario, contracts), the offline Audio Doctor (magnitude/frequency response, THD/THD+N, phase/group delay), and their Python sibling the Audio Quality Lab (tools/audio/quality-lab — null residual + alignment, LTAS log-spectral distance, spectral flux/centroid, HNR, Theil-Sen drift slope, Kaiser-sinc resampling, license-guarded corpus, regression-net ratchet). TRIGGER on AUTHORING work — "build/design an oscillator/filter/synth/effect", "add a DSP module", "what should the acceptance gate be", "how do I measure aliasing / anti-aliasing / alias floor", "null against a reference", "is this DSP correct", "choose a tolerance", "golden/regression corpus for audio", "measure drift or jitter", "A/B two renders" — AND on DEBUGGING work — "is there sound / no audio / I hear nothing", "does this filter/compressor/synth/delay produce the right signal", "prove the DSP / prove the contract", "measure the frequency response", "what's the THD / is it distorting", "what's the group delay / phase response / measured latency", "magnitude response curve", "render a test tone and assert", "audio regression", "64-frame works but 128 is silent", "sample-rate change pitch-shifted it", "describe what's in this buffer", "audio doctor", "compare before/after a DSP refactor". Reach for this BEFORE hand-rolling any FFT, null test, alias measurement, pitch tracker, or golden-render script — most of it already exists in one of the two lanes. Test/tool layer over HeadlessHost — deterministic, no audio device, no speakers. Off the realtime thread entirely.

日本語の概要は準備中です。原文の説明を表示しています。

Generous-Corp/pulp222026年10月10日 更新

Generous-Corp のスキルをすべて見る

このスキルの問題を報告する