本文へ移動
cccskills
無料GitHub で公開

trace-sql

SQL discipline for querying Pulp Perfetto traces (.pftrace) with trace_processor — idempotent CREATE OR REPLACE PERFETTO views, GLOB not LIKE, dur = -1 incomplete-slice handling, EXTRACT_ARG for span args, joining on stable utid/upid, SPAN_JOIN PARTITIONED, and the draft→validate→execute loop. Ships Pulp's CPU/UI trace stdlib plus closed GPU startup, health, and probe views. TRIGGER when writing or debugging SQL over a .pftrace, when `pulp trace query` returns wrong/empty rows, or when the trace-analysis skill needs a query primitive.

インストール方法を見る

含まれるファイル(12)

  • SKILL.md38.0 KB
  • pulp_dsp_node_cost.sql1.3 KB
  • pulp_frame_stage_cost.sql2.6 KB
  • pulp_frames_over_budget.sql1.3 KB
  • pulp_gpu_audio_blocks.sql30.6 KB
  • pulp_gpu_health_transitions.sql3.1 KB
  • pulp_gpu_probe_correlation.sql3.7 KB
  • pulp_gpu_startup_breakdown.sql8.1 KB
  • pulp_layout_vs_paint.sql1.4 KB
  • pulp_motion_join.sql1.2 KB
  • pulp_slowest_frames.sql1.1 KB
  • pulp_xruns.sql1.5 KB

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

trace-sql — querying Pulp traces with trace_processor

Pulp's tracing subsystem writes Perfetto .pftrace files; this skill is the SQL discipline for turning one into an answer. It is the query substrate under the trace-analysis harness and under pulp trace query / pulp trace <preset>. You are reading this because you are writing SQL over a .pftrace, a query came back empty or wrong, or the investigation harness needs a named primitive to query.

This skill is provider-agnostic by construction: it is a plain SKILL.md read identically by Claude Code and Codex from .agents/skills/. Nothing here depends on a specific agent, MCP server, or vendor.

Attribution. The SQL discipline below (the draft→validate→execute loop, GLOB-over-LIKE, incomplete-slice handling, stable-key joins, SPAN_JOIN guidance) is adapted from Google's android/skills perfetto-sql (Apache-2.0). The methodology is reused; the Android domain content (SurfaceFlinger, binder, ftrace, cpu governor) is not — every query here is authored against Pulp's own category taxonomy. See NOTICE.md.


The tool: trace_processor

Queries run through the trace_processor wrapper (a small Python launcher that resolves the precompiled trace_processor_shell binary). SQL is strictly offline: pulp trace query requires --trace FILE.pftrace and shells to trace_processor; it never forwards SQL to a live inspector:

# Offline SQL against a saved capture (no live inspector needed):
pulp trace query "SELECT name, dur FROM slice ORDER BY dur DESC LIMIT 20" \
  --trace /tmp/pulp-<ts>.pftrace

Pulp resolves the binary via $PULP_TRACE_PROCESSOR → the pinned Pulp-fetched build → $PATH; pulp trace doctor reports which. For zero-install, run pulp trace fetch once: it downloads the pinned trace_processor_shell (Perfetto v57.2, matching the tracing SDK), SHA-256-verified per platform, into $PULP_HOME — no manual install, no surprise download inside query. trace_processor is also a first-class entry in Pulp's tool registry, so pulp tool install trace-processor (and pulp tool info/doctor trace-processor) fetch the identical pinned artifact — same code path as pulp trace fetch. To drive trace_processor directly:

# Interactive SQL over a capture:
./trace_processor --query-string \
  "SELECT name, dur FROM slice WHERE category='dsp.node' ORDER BY dur DESC LIMIT 20" \
  /tmp/pulp-<ts>.pftrace

# Run a file of view definitions, then query (trace as the trailing positional):
./trace_processor -q .agents/skills/trace-sql/pulp_dsp_node_cost.sql /tmp/x.pftrace

Before using it, check it exists (skill-versioning rule): confirm the trace_processor wrapper is on PATH or in the project's tracing cache. If it is absent, run the explicit pulp trace fetch; query does not silently download or degrade to returning a trace path.

pulp trace start|stop --instance ID selects one exact broker-owned live instance for lifecycle control. The selector is transported outside lifecycle method parameters and is rejected by offline query and other non-lifecycle verbs; omission preserves the canonical safe no-selector behavior.

Output format. Offline query emits trace_processor's native table by default; --format table is the only accepted explicit format. The global --json flag wraps that table as an escaped output string for agents; --format json|csv is rejected rather than mislabeling native output.

The three named GPU analyses add a bounded subprocess boundary around this tool: at most 512 MiB of trace input, 120 seconds of wall time, and 4 MiB of combined stdout/stderr, with full process-tree termination on failure. Keep named-view output selective and bounded. If a query approaches those limits, reduce its columns/rows or shorten the capture; never weaken the runner to make an accidentally unbounded query appear successful. Free-form trace query does not yet inherit these same bounds and must be treated as a follow-up hardening surface rather than proof of the named analyzer's safety contract. The named analyzer opens without following the final symlink/reparse point and without blocking on a Unix FIFO/device replacement, verifies a regular-file handle, then snapshots those bytes into an exclusive private file (mode 0600 on Unix). It rejects replacement or growth using handle-derived filesystem identity (device/inode on Unix, volume/file ID on Windows) and passes only the snapshot path to trace_processor.


The Pulp trace-stdlib (query named primitives, not re-derived SQL)

Nine authored CREATE OR REPLACE PERFETTO definitions ship next to this skill. Load the ones you need, then SELECT from them — do not re-derive the joins by hand each time. Each .sql file carries a header comment explaining its shape.

View / functionAnswersL0 preset
pulp_slowest_framesframe spans, slowest firstpulp trace slowest-frames
pulp_dsp_node_costper-node DSP cost (count/total/mean/max)pulp trace dsp-hotspots
pulp_frames_over_budgetframes past the vsync budget (fn takes a budget)--preset frames-over-budget
pulp_xrunsxrun / deadline-miss instant eventspulp trace xruns
pulp_layout_vs_paintframe-pipeline cost split, one row per stagepulp trace layout-vs-paint
pulp_frame_stage_cost(name)per frame: stage SELF time (layout/canvas/js/text/state/render), whole-surface repaint requests, layout passestools/scripts/trace_frame_cost.py
pulp_motion_joinframes joined to their motion trace_id--preset motion-join
pulp_gpu_startup_breakdownranked startup GPU/render stagespulp trace gpu-startup
pulp_gpu_health_transitionshealth/device-loss evidencepulp trace gpu-health
pulp_gpu_probe_correlationprobe/readback evidence correlationpulp trace gpu-probe

pulp_gpu_health_candidates and pulp_gpu_probe_candidates ship inside those last two files. They are the pre-correlation candidate sets the closed views select from, published so a consumer can tell a refused cohort from an absent one rather than re-deriving the join. They answer no preset of their own.

A3 product spans supply low-cardinality debug annotations such as debug.gpu_evidence_id and debug.trace_evidence_id; the C++ trace macro call uses the unprefixed annotation name and Perfetto exposes it under debug.*. Keep SQL joins on the exact GPU ID and stable process instance. The named /trace gpu-startup|gpu-health|gpu-probe --trace FILE surface must resolve to these checked-in views through the same installed analyzer used by A2T. Do not replace a missing named result with one-off SQL in a terminal receipt.

The producer behind PULP_A3_CAMPAIGN_PRODUCER owns the raw trace capture for its real role and must preserve the health response's GPU and trace evidence IDs. gpu_first_visible_a3_external_adapter.py pins that producer and passes its digest-bound artifacts onward; it is not another analyzer and cannot turn an empty named view, a different process cohort, or hand-written SQL into a passing campaign analysis. The checked-in role producers additionally pin the external lifecycle driver and checked-in source-bound analyzer wrapper and retain the closed driver request/receipt. The wrapper prepares one exact analyzer in a fresh, config-free Cargo home/target, strips ambient Cargo/Rust runners and flags, and retains toolchain/source/output digests. The driver supplies the trace, not a trusted analysis sidecar; the producer runs the named query plus an invalid- trace negative and accepts only the exact health evidence ID, recorded host UPID, and PID that answered the producer's live-host nonce challenge before the shared verifier consumes the derived analysis. Its structural unverified result is distinct from the campaign budget verdict.

Do not let the offline A2T no-producer classification hide the real A3 product health-transition spans. Terminal acceptance also consumes the derived pre-change/compile-out/compiled-in-idle/active overhead receipt from gpu_first_visible_a3_trace_producer_overhead.py. Each active sample is replayed from its session metadata and real producer span with exact process-start, binary, session-config, xrun, and audio-thread facts. Arbitrary trace bytes are rejected, and the final active binary must equal one measured role product.

The four-state collector accepts production binary Perfetto only and replays it through Pulp's exact v57.2 platform SHA pin. Chrome JSON is planted-fixture-only. Reject unfinished slices, loss/no-flush stats, foreign UPIDs, reused session challenges, mismatched host PIDs, xruns, or producer spans on declared audio TIDs. Active queries must find gpu_health_transition_first_visible, require gpu_acquire → gpu_submit → gpu_present per sample, and count the complete b4ba exact 20-signature state/render/js inventory. A zero count for one of the other 17 signatures is reported not-covered, never zero-cost; absence of a mandatory stage is missing evidence. Each row also has a source-bound state_build_driver: the collector exports the exact revision, rebuilds under a default-deny/no-network sandbox, and requires rebuilt executable bytes plus the tracing sentinel state to match measurement. The retained source/build/toolchain artifacts are part of offline replay; SQL truth cannot compensate for missing product provenance.

Closed GPU cohort boundary. The named GPU analyses do not currently accept an evidence-ID selector, so the SQL must not flatten unrelated runs. Startup selects the earliest valid render-frame lifecycle carrying frame_index = 0 and returns that lifecycle's unindexed/frame-zero cold work separately from its later indexed steady-state rows; the single-ID fallback exists only for legacy traces with no indexed render frame. Startup joins slice → thread_track on the stable utid and intersects overlapping thread_state rows when scheduler evidence exists; CPU/non-running attribution is exposed only when those intervals cover the complete slice. Partial or absent coverage remains NULL rather than becoming a blocking claim. Unindexed work is cold only when a correlated frame-zero anchor proves that it began before first-frame completion; later unindexed work is unknown. Tooling-owned correlation rows are selected from candidates that carry an evidence ID; generic Dawn/Skia backend spans may remain untagged because they are not allowed to supply the cohort. The selected rows return no result when their tagged candidates carry multiple IDs or lack one valid 32-lowercase-hex ID. The closed analyzer interprets an empty mixed-ID result as unavailable. A future multi-run UX should add an explicit evidence-ID selector rather than weakening this singleton boundary. Category discovery follows the same boundary: join slices through thread_track.utid to the stable thread.upid, and accept categories only when the question rows and category rows resolve to one evidence ID and one UPID/PID process instance. Never union global trace categories into a scoped answer. Probe verdict SQL also owns a closed causal-failure diagnostic set; cpu_oracle_mismatch and magnitude_dispatch_failed remain failures even if the same row inconsistently reports health_state=healthy. Any tooling-owned gpu_probe* or gpu_readback* candidate without an evidence ID invalidates the whole probe cohort; a healthy tagged row cannot hide it. Generic untagged backend work such as gpu_submit* remains allowed but cannot supply the cohort. Do not generalize the diagnostic rule to every nonempty diagnostic because healthy diagnostics are valid.

A refusal names itself. Both evidence-gated views answer an uncorrelatable capture and a capture holding none of their work the same way: with no rows. So each publishes its candidate set as its own view — pulp_gpu_probe_candidates and pulp_gpu_health_candidates — and the analyzer counts the question's own candidates before it describes an empty answer. Candidates present with no answer is invalid-evidence-correlation; no candidates at all is missing-question-category. The probe view's is_tooling_owned column carries the gpu_probe* / gpu_readback* classification that both the cohort rejection and that diagnosis read, so a refusal cannot be described by a second, drifting copy of the rule. Never restate either predicate in a consumer.

One definition, three surfaces. The L0 CLI preset names map 1:1 onto these views: slowest-frames → pulp_slowest_frames, xruns → pulp_xruns, dsp-hotspots → pulp_dsp_node_cost, layout-vs-paint → pulp_layout_vs_paint. The same view is the preset verb, the agent's query primitive, and the human's SELECT target. When you extend the stdlib, keep that mapping intact.

Load and query in one shot:

./trace_processor \
  -q .agents/skills/trace-sql/pulp_slowest_frames.sql \
  --query-string "SELECT * FROM pulp_slowest_frames LIMIT 10" \
  /tmp/x.pftrace

The category vocabulary the views key off is the fixed taxonomy: dsp, dsp.node, render, layout, canvas, text, js, gpu, state, io.


The draft → validate → execute loop

Do not fire a hand-written query blind against a large trace. Iterate:

  1. Draft the query from the question and the category taxonomy.
  2. Validate cheaply — append LIMIT 5 (or wrap in SELECT COUNT(*)) and run it. Confirm the columns, the category filter, and that rows come back at all. Fix schema/typo/empty-result problems here.
  3. Execute the full query only once the shape is proven.

Cap validation at ≤ 3 iterations. If three drafts still return nothing, the problem is usually not the SQL — it is the capture (wrong categories traced, ring overflowed to empty, or the span simply was not emitted). Switch to diagnosing the trace, not the query.


SQL discipline (the rules that bite)

Idempotency — always CREATE OR REPLACE PERFETTO. Views/functions must be re-runnable in the same trace_processor session without a "already exists" error. Every stdlib file uses CREATE OR REPLACE PERFETTO VIEW / ... FUNCTION. Never a bare CREATE VIEW.

GLOB, not LIKE. PerfettoSQL GLOB is case-sensitive and uses * / ? — the behavior trace tooling expects (name GLOB 'xrun*'). LIKE is case-insensitive and uses % / _; reaching for it silently changes matching semantics and folds unrelated spans in. Use GLOB for name patterns.

dur = -1 means incomplete, not instant. A slice still open when the ring flushed (or an unterminated PULP_TRACE_BEGIN) has dur = -1. It is not a zero-length event. Every duration query filters WHERE dur >= 0 (or dur != -1). A negative duration leaking into a SUM/ORDER BY corrupts the result. Instant events genuinely have dur = 0 — key those off the name, not the duration. Xruns are not the only ones: GPU diagnostics arrive as dur = 0 slices named gpu.diagnostic on category gpu, carrying their severity and message as debug. args (EXTRACT_ARG(arg_set_id, 'debug.severity')).

EXTRACT_ARG for span arguments — mind the debug. prefix. Typed args (frame index, block index, motion.trace_id, sample position) live in the arg set, not as columns. Pulp emits them as TRACE_EVENT debug annotations, so trace_processor keys them debug.<name>, not bare: EXTRACT_ARG(s.arg_set_id, 'debug.frame_index'), not 'frame_index'. A missing arg (or a wrong/bare key) yields NULL silently — this is the number-one reason a preset returns empty rows. Confirm the real key with SELECT DISTINCT key FROM args against the trace before trusting a query, and filter IS NOT NULL when the arg is a join key.

Join on stable identity — utid / upid, never tid / pid. OS thread and process ids are recycled within a trace; Perfetto's utid (unique thread id) and upid (unique process id) are stable for the whole capture. Join slice → thread_track → thread on utid, and reach the process via thread.upid. Joining on raw tid/pid will cross-link two threads that reused an id.

SPAN_JOIN needs PARTITIONED. To intersect two span tables on the time axis (e.g. correlate dsp slices with a concurrent render window), use SPAN_JOIN and partition by the stable key so spans are only joined within the same thread/track:

CREATE VIRTUAL TABLE dsp_x_render USING SPAN_JOIN(
  dsp_spans   PARTITIONED utid,
  render_spans PARTITIONED utid
);

Omitting PARTITIONED cartesian-joins across unrelated threads and both explodes the row count and produces meaningless overlaps.

Percentiles. SQLite has no percentile builtin. The stdlib views expose mean_ms alongside max_ms precisely so the mean-vs-max gap flags a per-block spiker without a percentile. When the trace-analysis harness needs true p95/p99 tail latency, compute it over the raw slice rows with a window/rank approach (NTILE, or ROW_NUMBER() ... ORDER BY dur over the node's slices) rather than trusting the mean.


Common query shapes

-- Slowest slices in a category (the flamegraph's fattest bars):
SELECT name, dur/1e6 AS dur_ms
FROM slice
WHERE category = 'dsp.node' AND dur >= 0
ORDER BY dur DESC LIMIT 20;

-- Which thread did a span run on:
SELECT s.name, thread.name AS thread, s.dur/1e6 AS dur_ms
FROM slice s
JOIN thread_track tt ON s.track_id = tt.id
JOIN thread ON thread.utid = tt.utid
WHERE s.category = 'text' AND s.dur >= 0
ORDER BY s.dur DESC LIMIT 20;

-- Count of a span by name (is TextShaper::prepare firing every frame?):
SELECT name, COUNT(*) AS n, SUM(dur)/1e6 AS total_ms
FROM slice
WHERE category = 'text' AND dur >= 0
GROUP BY name ORDER BY n DESC;

-- Steady-state vs whole-run average per node (drop the single cold-start
-- outlier). A per-node average is often dominated by the FIRST block, where
-- the first process() call warms caches / touches fresh pages — a one-time
-- cost, not a per-block cost. Subtracting the per-node max isolates it:
SELECT name,
       ROUND(AVG(dur)/1e3, 2)                          AS avg_all_us,
       ROUND((SUM(dur) - MAX(dur)) / 1e3 / (COUNT(*) - 1), 2) AS avg_steady_us
FROM slice
WHERE category = 'dsp.node' AND dur >= 0
GROUP BY name ORDER BY avg_all_us DESC;

Gotchas

  • Empty result ≠ zero cost. An empty query often means the category was not traced, or the ring overflowed and the trace is silently truncated. Check SELECT DISTINCT category FROM slice before concluding "nothing was slow."
  • Wall time, not CPU time. slice.dur is wall-clock span duration. A long slice may be blocked, not computing — the trace-analysis skill's wall-time-vs-CPU-time rule (join to thread_state) decides which.
  • Category is a string column. WHERE category = 'dsp' does not match 'dsp.node' — they are distinct taxonomy entries. Use GLOB 'dsp*' only when you deliberately want both.
  • Dynamic span names don't aggregate under GROUP BY name. Most spans use a static name (the default), so GROUP BY name collapses them into one row per callsite — that is what makes "is TextShaper::prepare firing every frame?" answerable. A span emitted through PULP_TRACE_SCOPE_DYNAMIC(cat, expr) carries a runtime-computed name (a node id, a parameter name), so each distinct value is its own slice.name and a GROUP BY name fragments instead of aggregating. If a per-name rollup looks unexpectedly scattered, check whether the callsite is dynamic (SELECT name, COUNT(*) FROM slice WHERE category=... GROUP BY name will show the high-cardinality spread); group by category or a name prefix (substr(name, 1, instr(name,'_'))) instead when you need the aggregate.
  • A self-join on a nullable key needs IS, not =. pulp_gpu_startup_breakdown.sql correlates each row to its cold-frame anchor by evidence id. A capture carrying no instrumentation at all is admitted as one untagged cohort whose evidence_id is NULL, and anchor.evidence_id = c.evidence_id never matches NULL against NULL — the correlation returns no anchor, every row reports a NULL cold_frame_end_ts, and the classification degrades silently instead of failing. Use SQLite's null-safe IS (anchor.evidence_id IS c.evidence_id) on any join key a cohort may legitimately leave NULL.
  • Gate a relaxation on the ABSENCE of the strict population, never per row — and scope that population to the whole trace, not to the question's own candidate set. The untagged cohort is admitted by WHERE NOT EXISTS (SELECT 1 FROM trace_evidence), where trace_evidence scans args for debug.gpu_evidence_id / args.debug.gpu_evidence_id across the entire capture. A per-row OR evidence_id IS NULL reads the same in the happy case and admits a capture that carried evidence and then lost some of it. The subtler miss is gating on identified_candidates: the startup candidate predicate excludes gpu_probe*, gpu_readback* and gpu_health_transition, so a capture whose probes are tagged while its startup spans are not has an empty identified_candidates and reads as uninstrumented — gpu-probe answers that same file pass with a real evidence id and a bound category_scope. Whatever "the strict population" means for your question, measure it where the producer writes it.
  • A cohort admitted without an identity needs its own boundary. Absence of evidence is necessary but not sufficient: untagged rows carry no id to group by, so nothing separates two lifecycles the way one distinct evidence_id per lifecycle does on the tagged path. admissible_untagged_cohort therefore also requires frame_zero_anchor_count <= 1 and process_count <= 1. Without the anchor bound, one single-process capture holding two frame-zero anchors merges the second lifecycle's setup into the first lifecycle's cold answer and ranks it dominant — and a process-scope guard alone (COUNT(DISTINCT COALESCE(upid, -1)) = 1) does not catch it, because both lifecycles are in one process.
  • These .sql files are compiled into the Rust tool by include_str!. experimental/pulp-rs/src/cmd/trace_gpu_analysis.rs embeds pulp_gpu_startup_breakdown.sql, pulp_gpu_health_transitions.sql, and pulp_gpu_probe_correlation.sql at build time. Two consequences: an uncommitted edit to one of them is exactly what cargo test exercises (no commit needed before measuring), and each question reads one file — a failing gpu-probe test cannot have been caused by an edit to the startup view.

Files this skill covers

  • .agents/skills/trace-sql/pulp_slowest_frames.sql
  • .agents/skills/trace-sql/pulp_dsp_node_cost.sql
  • .agents/skills/trace-sql/pulp_frames_over_budget.sql
  • .agents/skills/trace-sql/pulp_xruns.sql
  • .agents/skills/trace-sql/pulp_layout_vs_paint.sql
  • .agents/skills/trace-sql/pulp_motion_join.sql
  • .agents/skills/trace-sql/pulp_gpu_startup_breakdown.sql
  • .agents/skills/trace-sql/pulp_gpu_health_transitions.sql
  • .agents/skills/trace-sql/pulp_gpu_probe_correlation.sql
  • .agents/skills/trace-sql/pulp_gpu_audio_blocks.sql — independent admission/terminal and eligibility/delivery identities; see docs/guides/gpu-audio-tracing.md for its explicit-load validator.
  • core/runtime/include/pulp/runtime/trace.hpp — the macro surface + category taxonomy
  • docs/guides/tracing.md — the guide, tiers, and worked use cases

Match trace_processor to the SDK's Perfetto pin

PulpTracing.cmake pins the Perfetto SDK (PULP_PERFETTO_VERSION, v57.2 at time of writing). Query with the same version — pulp trace fetch, or curl -sSL -o trace_processor https://get.perfetto.dev/trace_processor which self-reports its version on --version.

The capture producer and query tool are separate compatibility surfaces. An SDK-built plug-in must link the installed SDK's exported Perfetto support from the same PULP_TRACING=ON configuration; a consumer-side define cannot add targets or headers omitted at SDK install time. Once a capture exists, query it with the trace processor matching that SDK pin. When diagnosis starts with “no rows,” first establish whether a .pftrace was actually flushed; do not debug SQL against a missing producer.

Cost that lives BETWEEN slices

pulp_layout_vs_paint and friends aggregate slice durations, so they cannot see a stall that happens inside a parent span but outside every child — the classic case being a blocking swapchain acquire. Two queries catch it:

-- 1. unaccounted time inside a frame: children summing to far less than the parent
SELECT s.id, s.dur/1e6 AS frame_ms,
       (SELECT SUM(c.dur)/1e6 FROM slice c WHERE c.parent_id = s.id) AS children_ms
FROM slice s WHERE s.name = 'frame' AND s.dur >= 0
ORDER BY (s.dur - IFNULL((SELECT SUM(c.dur) FROM slice c WHERE c.parent_id = s.id),0)) DESC
LIMIT 20;

-- 2. inter-frame gaps: idle or blocked BETWEEN frames
WITH f AS (SELECT ts, dur, LEAD(ts) OVER (ORDER BY ts) AS nts
           FROM slice WHERE name = 'frame' AND dur >= 0)
SELECT COUNT(*) AS gaps, AVG((nts-(ts+dur))/1e6) AS mean_gap_ms,
       MAX((nts-(ts+dur))/1e6) AS max_gap_ms
FROM f WHERE nts IS NOT NULL;

Query 1 is what identified the Windows knob-drag regression: frames of 19-45 ms whose children summed to ~2 ms. Small gaps plus large unaccounted time means the thread is blocked inside the frame, not idle between frames.

An empty capture may be a missing attachment, not a bad query (WAH-4)

Before reaching for trace_processor, confirm the session actually recorded. Tracing attach/detach was wired into VST3 only until WAH-4; a capture of a CLAP, AU v2, AU v3, AAX or Standalone session produced an empty .pftrace while every command looked correct. All six formats now attach via runtime::ScopedTracingAttachment.

Two other capture-shaping facts worth knowing before you blame a query:

  • The trace is written by the FINAL detach. A leaked attachment means it is never written at all. Let the host finish unloading the plug-in rather than killing the process.

  • PULP_TRACE_SECONDS timeouts are tagged with a session generation. They used to be untagged, so closing and reopening an editor inside the window let the FIRST session's timer stop the SECOND one — a capture that looks mysteriously truncated mid-gesture. A stale timer is now a no-op.

  • Zero rows everywhere can mean the binary was never built with tracing. PULP_TRACING is OFF by default, and a build without it emits no spans at all while every command still succeeds — the emptiest possible capture from the healthiest-looking run. This is the one absence to rule out first, because it is indistinguishable from a bad query by inspection of the query. pulp status prints a Tracing: line for the current checkout, and pulp build --trace produces a build that can emit. A prefix or directory whose name contains trace proves nothing; only the cache does.

  • Zero rows for frame/gpu_* can mean the host never emitted them. A query over render spans returning nothing is not automatically a bad query or a bad capture: an editor that is neither scripted nor declares requires_gpu_host() runs on CPU raster and emits none of them, and host-level frame/paint exist only on the Windows plug-in editor and the macOS standalone host. Check the shape of what you DID capture before rewriting SQL —

    select name, count(*) from slice group by name order by 2 desc;
    

    A result of only layout_children + wm_mousemove is the signature. The trace-analysis skill has the coverage matrix and the [plugin-gpu-host] … mode= line that names the host you actually got.

A zero-row query on a WRAPPED ring is not a finding

The stats table is the only thing that distinguishes "the span never fired" from "the capture is unreadable", and the second case is the one that looks clean. When the in-process ring wraps, the interned string table at the head of the sequence is overwritten and every later packet on that sequence is skipped; trace_processor opens the file, reports no error, and answers every query with zero rows. The file on disk is full-size, so nothing about it looks wrong.

Run this before quoting a number out of any capture:

select name, value from stats
where value != 0 and name in (
  'traced_buf_write_wrap_count',
  'traced_buf_bytes_written',
  'traced_buf_buffer_size',
  'traced_buf_bytes_overwritten',
  'traced_buf_incremental_sequences_dropped',
  'packet_skipped_seq_needs_incremental_state_invalid');

A non-zero traced_buf_write_wrap_count, or any packet_skipped_seq_needs_incremental_state_invalid, condemns the trace. Do not analyse it and do not report an absence from it — re-capture with PULP_TRACE_RING_KB raised (KB; default 80 MB; a UI capture carrying js_native spans needs ≥ 262144, and 60 s with audio playing ≈ 1572864).

Note the asymmetry: a clean stats table proves only that nothing overflowed, not that anything recorded. Pair it with a positive control whose count MUST be non-zero for the workload you captured, e.g. select count(*) from slice where name='dom_event_evaluate' for a pointer-driven UI capture.

GPU render time is now OPT-IN (WAH-13)

SkiaSurface::gpu_render_timing_available() reporting false is no longer evidence of an adapter that lacks timestamp-query. Timestamps are requested only when the host asks, via PluginViewHost::Options::enable_gpu_timing (default OFF), rather than whenever the adapter advertises the feature.

That default is deliberate and worth understanding before you "fix" it: Dawn gates writeTimestamp behind the allow_unsafe_apis toggle on every backend, so requesting the feature forces that toggle on — and it applies to the DEVICE, not to the diagnostic. Ordinary rendering was silently running with relaxed validation on every machine whose adapter happened to offer timestamps.

If you need per-recording GPU time in a capture, enable it explicitly on the host's Options. If a trace shows no gpu_render_time, check that flag before suspecting the adapter.

A3 v2 terminal evidence is analyzer-derived rather than sidecar-attested. The pinned replay must bind the exact trace digest, role/campaign/instance/build identity, GPU and trace evidence IDs, process PID/UPID, the closed category set, zero drops, and a completed flush.

Self time is the query that finds an opaque parent

A slice's dur includes everything it called, so a parent that dominates a duration ranking tells you nothing about where the time went. Rank by self time — total minus the sum of direct children — and an unattributed block names itself:

SELECT s.name, COUNT(*) AS n,
       SUM(s.dur)/1e6 AS total_ms,
       (SUM(s.dur) - IFNULL(SUM((SELECT SUM(c.dur) FROM slice c
                                 WHERE c.parent_id = s.id AND c.dur >= 0)), 0))/1e6 AS self_ms
FROM slice s WHERE s.dur >= 0
GROUP BY s.name ORDER BY self_ms DESC LIMIT 20;

Beware the shape of the subquery: summing all descendants instead of direct children double-counts nested time and drives self time negative. parent_id (not a recursive descent) is the correct join, and a negative self_ms in the output means you got it wrong, not that the trace is broken.

This is the query that showed dom_event_evaluate holding 2119.8 ms of self time out of 2213.6 ms total — 95.8% of a script event handler outside every span the tree emitted.

Querying the JS bridge spans

js_native slices (one per JS→C++ bridge call, tracing builds only) and script-authored pulpTrace spans both land on the js category. The slice name is the constant js_native; the bridge function is its debug.fn arg. Attribute a handler's native half with:

SELECT EXTRACT_ARG(arg_set_id, 'debug.fn') AS fn,
       COUNT(*) AS calls, SUM(dur)/1e6 AS ms, MAX(dur)/1e6 AS max_ms
FROM slice WHERE category = 'js' AND name = 'js_native' AND dur >= 0
GROUP BY fn ORDER BY ms DESC LIMIT 25;

An older naming put the function in the slice name (js_native:<fn>); a GLOB 'js_native:*' filter against a current trace matches nothing and reads as "no bridge calls". Pair it with SELECT COUNT(*) FROM slice WHERE name = 'js_native' before believing a zero.

Check for corrupted parentage before trusting any js aggregate. An unbalanced __traceBegin__ re-parents later slices under a span that never closed; the runtime force-closes it and marks the trace. This must return zero rows, and if it does not, the nesting is untrustworthy from that timestamp on:

SELECT COUNT(*) FROM slice WHERE name = 'js_trace_force_closed_unbalanced_scope';

Its positive control is the counter track js_trace_unbalanced_scopes, which carries the leaked depth as a value rather than a count of incidents.

Check the category is populated before trusting a zero

trace.hpp declares ten categories and not all of them emit. A WHERE category GLOB '<name>' filter over one that emits from nowhere returns no rows — which looks identical to "that subsystem did no work".

Which categories are empty CHANGES as code is instrumented, so measure rather than trust a list:

for c in dsp dsp.node render layout canvas text js gpu state io; do
  printf '%-10s %s\n' "$c" \
    "$(git grep -oE "PULP_TRACE_[A-Z_]+\(\s*\"$c\"" origin/main -- core inspect | wc -l)"
done

Then pair any per-category question with a control that must return rows:

-- the finding
SELECT COUNT(*) AS n FROM slice WHERE category GLOB '<target>';
-- the control: a category the sweep above shows is populated
SELECT COUNT(*) AS n FROM slice WHERE category GLOB 'gpu';

Control non-zero and finding zero means the instrumentation is absent, not the work. Report the gap; do not report a timing verdict.

The same discipline is written up at more length in the trace-analysis skill under "A declared category is not a populated one" — keep the two in step if either changes.

WaveNet admission, delivery, and unresolved ownership

For pulp_gpu_audio_blocks.sql, correlate (upid, engine_id, generation, sequence). Match actual drained admissions to exactly one GPU terminal, with no duplicate or orphan identities. Callback eligibility/delivery is separate: fallback can occur without GPU admission, and a published GPU result need not have been delivered. Use gpu_reason for the GPU result and delivery_reason for the callback decision. A known submission failure remains submission_rejected after teardown; a stale result retains its recovery cause, including input_saturated, device_lost, teardown, or invalid_callback. A deadline_exceeded callback reason establishes that no usable due result was available, not why it was unavailable.

callback_ingress_ns is observed callback entry, not the intended schedule. ingress_to_worker_ns measures admission delay after that observation; worker_to_observed_ns includes provider work and completion servicing. Missing scheduled/encode/submit/GPU timestamps stay unavailable. Do not label either span GPU execution time or infer scheduler latency from it alone.

The audio block view also projects nullable batch_id, microbatch_size, model-family/hash, provider identity hashes, and predictor/version/calibration fields from terminal and delivery records. Use the session claim_* flags when making a batching, model/provider identity, or predictor claim: the view emits explicit missing_* issues for any incomplete claimed block and flags terminal/delivery one-sided omissions or mismatches. Historical schema-2 rows with no claim flags remain compatible; missing metadata is unavailable, never zero. deadline_margin_ns is signed, so negative values are valid.

gpu.audio.ownership is a non-sequence event. If destruction cannot establish physical release, it records physical_release_complete=false and unresolved_channel_count, the number of channel releases that returned false. It does not assert a GPU terminal or retirement. The SQL must report unresolved_physical_ownership and reject complete-lifecycle acceptance even if all block identities otherwise close. A failed release followed by a successful retry must not produce this final unresolved-ownership disclosure. Require the canonical validator's zero-loss checks as well as identity closure; aggregate counters or an empty trace cannot substitute for actual records.

See docs/guides/gpu-audio-wavenet-trace.md for the capture contract and current physical NAM measurement boundaries.

Traces from an out-of-process host, and prewarm proof

A traced build reads ~/.config/pulp/trace-autostart when PULP_TRACE_PATH is not in its environment (AUHostingService inherits none). A PULP_TRACE_PATH ending in / writes one file per process, <process>-<pid>.pftrace, so a host and its service never overwrite each other — load the service's file, not the host's, for the editor's spans. The background prewarm runs on its own thread, so key on the thread when proving it: script_precompile / scripted_ui_prewarm on the worker, and only script_bytecode_read (no script_compile) for the same script inside editor_first_frame on the main thread. Join slice → thread_track → thread (stable utid) rather than assuming one track per process.

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

同じリポジトリのスキル

概要と使いどころ

aax

無料

Optional AAX support for Pulp, including developer-supplied Avid SDK setup, CMake enablement, DigiShell/AAX Validator workflows, and local AAX builds on macOS or Windows.

日本語の概要は準備中です。原文の説明を表示しています。

Generous-Corp/pulp222026年10月10日 更新

Configure, implement, and test Pulp's optional desktop Ableton Link tempo-sync adapter while preserving the developer-supplied SDK, licensing, realtime, latency-compensation, and no-install boundaries.

日本語の概要は準備中です。原文の説明を表示しています。

Generous-Corp/pulp222026年10月10日 更新

Maintain Pulp's installed design-time agent capability manifest and public-surface ledger. Use when adding, removing, renaming, or materially changing public audio, MIDI, signal, timebase, or sequence APIs; registering a new algorithm for generators; changing capability support or deprecation state; or repairing agent-capabilities freshness, schema, fingerprint, tombstone, or installed-SDK tests.

日本語の概要は準備中です。原文の説明を表示しています。

Generous-Corp/pulp222026年10月10日 更新

android

無料

Android platform development for Pulp — NDK cross-compilation, Oboe audio, Dawn/Skia GPU rendering, JNI bridge, touch interaction, emulator workflows, and end-to-end smoke validation. Covers build, deploy, debug, and the gotchas discovered during bringup.

日本語の概要は準備中です。原文の説明を表示しています。

Generous-Corp/pulp222026年10月10日 更新

ara

無料

Optional ARA support for Pulp, including developer-supplied ARA SDK setup, CMake enablement, adapter companion APIs, validation, and ARA-aware plugin implementation guidance.

日本語の概要は準備中です。原文の説明を表示しています。

Generous-Corp/pulp222026年10月10日 更新

The measurement surface for ALL Pulp DSP and audio-pipeline work — read it BEFORE writing or gating DSP, not only when something already sounds wrong. Covers the C++ harness (signal generators, metrics, assertions, RenderScenario, contracts), the offline Audio Doctor (magnitude/frequency response, THD/THD+N, phase/group delay), and their Python sibling the Audio Quality Lab (tools/audio/quality-lab — null residual + alignment, LTAS log-spectral distance, spectral flux/centroid, HNR, Theil-Sen drift slope, Kaiser-sinc resampling, license-guarded corpus, regression-net ratchet). TRIGGER on AUTHORING work — "build/design an oscillator/filter/synth/effect", "add a DSP module", "what should the acceptance gate be", "how do I measure aliasing / anti-aliasing / alias floor", "null against a reference", "is this DSP correct", "choose a tolerance", "golden/regression corpus for audio", "measure drift or jitter", "A/B two renders" — AND on DEBUGGING work — "is there sound / no audio / I hear nothing", "does this filter/compressor/synth/delay produce the right signal", "prove the DSP / prove the contract", "measure the frequency response", "what's the THD / is it distorting", "what's the group delay / phase response / measured latency", "magnitude response curve", "render a test tone and assert", "audio regression", "64-frame works but 128 is silent", "sample-rate change pitch-shifted it", "describe what's in this buffer", "audio doctor", "compare before/after a DSP refactor". Reach for this BEFORE hand-rolling any FFT, null test, alias measurement, pitch tracker, or golden-render script — most of it already exists in one of the two lanes. Test/tool layer over HeadlessHost — deterministic, no audio device, no speakers. Off the realtime thread entirely.

日本語の概要は準備中です。原文の説明を表示しています。

Generous-Corp/pulp222026年10月10日 更新

Generous-Corp のスキルをすべて見る

このスキルの問題を報告する