本文へ移動
cccskills
無料GitHub で公開

proxy-first-eval

Judge whether a change made things better with demonstrable proxy measures tied to the mechanism it targets, not wall-clock time. Load before choosing acceptance measures for any before/after or A/B claim — CI and fleet speed-ups, build speed, DSP/audio quality, render fidelity. Each proxy names its data source, detection floor and sample size, and every zero is paired with a control on the same instrument.

インストール方法を見る

含まれるファイル(1)

  • SKILL.md10.9 KB

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

Proxy-first evaluation

Load this before you pick how to measure a change, not after the numbers are in. Any claim of the shape "X is now better / faster / closer" is in scope: a CI routing fix, a fleet capacity change, a build-speed tweak, a DSP rewrite, an importer or renderer fix.

The rule

Pick the proxy from the mechanism the change targets. Ask: what does this change physically alter? Measure that, at the point where it happens, in a form someone else can re-count from a log line, a label or an annotation.

Wall-clock time is almost never that. It mostly tracks load, queue depth, neighbouring jobs and noise. A change can make things better while wall time rises (the fleet got busier) or leave them unchanged while wall time falls (a quiet evening). Wall time is context, never the verdict — report it, label it load-dependent, and do not let it decide.

Proxies by mechanism

CI / fleet

Change targetsProxySource
Placement (routing a job to the right runner class)share of jobs that landed on a runner of the intended classcompleted job runner_name / runner_group_name / labels, not the workflow's runs-on:
Starvationshare of jobs cancelled before any runner was assignedjobs with conclusion=cancelled and empty runner_name
Queue waitwait per job ahead in the queue, not raw waitcreated_at → started_at, divided by jobs queued ahead at created_at
Gate costrequired-gate runs (or minutes) per merged PRshipyard metrics gate-cost
Merge-queue churnmerge-queue attempts per merged PRmerge_group runs / PRs merged in the window
Queue ejectionsejections by cause (red check, timeout, wedge, neighbour failure)merge-queue events + the failing check-run's output.title
Build speedcompile units rebuilt, cache hit rate, blast radius of a headertools/scripts/build_speed_scorecard.py report --split <ISO-time>
Test-result reuse: benefittest-seconds skipped ÷ merge-group test-seconds, per group (ctest per-test durations, never elapsed wall time)ctest log lines / JUnit time per test; reuse_policy_replay.py score
Test-result reuse: safetyfalse skips (group tests that FAILED which the policy would have skipped)replay corpus outcome vs policy verdict; live: sampled re-run failures
Test-result reuse: flakesflake-skips (skipped tests whose fail was retried to pass or exonerated)attempts > 1 in JUnit; exoneration shadow annotation
Receipt supplyreceipts issued ÷ eligible PR headsissuer notice / artifact listing vs full-suite green heads
Receipt usereuse ÷ evaluated groupsshipyard-receipt-decision/v1 annotations, by verdict and reason

Normalise every count for volume: a rate per job, per PR or per merge, never a raw total across windows of different traffic.

A reuse policy (anything that lets a merge group skip tests an earlier run proved) ships only when tools/scripts/reuse_policy_replay.py score reads 0 false skips over the history window; its benefit is reported beside that, never instead of it. Controls on the same corpus: the none policy reads 0% benefit, whole-receipt on a tree-identical pair reads 100%, and the score --scenarios fixtures include a synthetic failing record that must read 1 false skip. A policy whose coverage is small is "insufficient sample", not safe.

DSP / audio

Change targetsProxyTool
Correctness vs referencenull residual with alignment (dB)assert_null_near, quality-lab compare
Aliasing / distortiontone residual by least-squares projection, THD/THD+Ntone_residual_db() prior art, Audio Doctor
Perceptual artifactsdetector counts with timestamps (transient smear, dulling, metallic HF, graininess)pulp tool run audio-quality-lab -- compare (/audio-compare)
Filter shapemagnitude response at named frequenciessignal::frequency_response, Audio Doctor

A listening impression is a pointer to where to measure, not a verdict. See the audio-harness skill for the lanes and the window-floor traps (Hann cannot see −100 dBc; the default OversamplerT kind has ~7 dB alias rejection).

Realtime GPU audio

Change targetsProxySource
Lower transport latencyminimum lead blocks meeting the delivery target; report intrinsic processor latency and transport lead separatelyauthenticated gpu.audio.session plus terminal/delivery records, grouped by block size and sample rate
Batching or graph-wide schedulingGPU delivery rate, batch/microbatch size, queue depth, and callback deadline tail at each leadexact model/provider/engine/generation/sequence receipt rows
Execution-time predictionprediction error and calibrated deadline margin (p50/p95/p99/max), plus the lead selected from that marginpredictor version/calibration and predicted-vs-observed timing fields in the same receipt
CPU fallback qualityfallback rate, late/drop/miss count, callback p99.9/max, and CPU-oracle null residualterminal disposition, delivery disposition, callback timing, and matched CPU-only control

For a lead-reduction claim, the verdict is the smallest lead that meets the declared GPU-delivery target (normally at least 99% over a complete campaign) while preserving zero callback deadline misses through the exact CPU fallback contract. A faster kernel or lower wall time does not establish that verdict. Run the same workload at each slots × lead × batch cell, with a CPU-only control and planted saturation/late/drop controls on the same instrument. Keep provider, executable, model, host, Release build, sample rate, block size, contention state, and trial length fixed; retain raw records and state n per cell. Do not infer a TCN, LSTM, compact-SSM, or Mamba result from a WaveNet receipt.

Render / import fidelity

Change targetsProxyTool
Layoutper-node box deltas in pxlayout_parity.py
Material survivalproperties present in the envelopematerial_audit.mjs
A named regionper-region scorediff_against_reference_regions.py
Controls workdriven-control assertionsthe prove-before-showing skill

A whole-image similarity score is position-blind triage, not a fidelity verdict. Read each tool's Cannot see line in the CLAUDE.md tool registry before quoting its number.

Demonstrable means three numbers per proxy

  1. Data source — the exact log line, label, annotation or API field, so a reviewer can re-count it.
  2. Detection floor — the smallest effect this instrument can see. Prove it with a negative control (run it on a case with the defect removed and show the reading collapses), don't derive it.
  3. Sample size — n per side. Small n (a handful of runs, one merge window, one render) is "insufficient sample", not a verdict in either direction.

Controls and instrument traps

  • Pair every zero with a control on the same instrument and target that must return non-zero. If the control is also zero, the instrument is broken; report nothing. Compare the control's count to what you expect, not just "non-zero".
  • Identical results across different filters means the filter is ignored. Example: actions/runs?workflow_id=... silently ignores the parameter and returns every workflow's runs; use actions/workflows/<file>/runs.
  • Do not grep whole job logs for a marker. Logs echo the step's own script, so the pattern matches its own source. Count annotations, ##[notice] / ##[error] lines, or check-run output instead.
  • actions/jobs/<id> handed a check-run id returns a coherent, wrong job. Use check-runs/<id> for the merge gate's own record.
  • Watch the failure shape, not only success. A job queued with no runner ever assigned, a merge-queue entry with no merge_group run, a test that SKIPs — all read as "nothing bad happened" to a success-only query.
  • Confirm the before and after measured the same thing: same workflow, same job name, same stimulus, same canvas size, same build type (Release).

Checklist

  • Named the mechanism the change targets, in one sentence.
  • Chose a proxy measured at that mechanism, normalised per job / PR / merge.
  • Wrote down the data source a reviewer can re-count.
  • Stated the detection floor, proven by a negative control.
  • Ran a positive control for every zero.
  • Recorded n per side; declared "insufficient sample" if it is small.
  • Checked the failure shape (no runner, no run, SKIP), not only success.
  • Reported wall time as load-dependent context only.

Report template

Mechanism:   <what the change physically alters>
Proxy:       <measure, normalised>            before -> after
Source:      <log line / label / API field / tool invocation>
Floor:       <smallest detectable effect; negative control used>
n:           <before n> / <after n>   (insufficient sample if < ...)
Controls:    <positive control for each zero, with its count>
Verdict:     better | worse | no change | insufficient sample
Context:     wall time <before -> after>, load-dependent, not the verdict

For a GPU-audio campaign, append:

Workload:    <model family/id/hash, provider/source/executable identities>
Geometry:    <sample rate, block size, slots, requested lead, batch size>
Latency:     intrinsic <samples> + transport lead <blocks> = reported <samples>
Prediction:  <version/calibration>, predicted vs observed p50/p95/p99/max, margin
Delivery:    GPU <count/rate>, CPU fallback <count/rate>, late/drop/miss <counts>
Control:     CPU-only + positive saturation/late/drop controls; n per cell

Tools

  • shipyard metrics gate-cost — gate minutes per merged PR, batch fullness, receipt reuse. shipyard metrics compare for before/after windows; treat its timing columns as context.
  • tools/scripts/build_speed_scorecard.py report --split <ISO-time> — build proxies split at the change.
  • tools/scripts/reuse_policy_replay.py collect|score — replays a test-result reuse policy over merge-queue history: benefit, false skips, flake-skips, coverage, and the named incident scenarios.
  • audio-harness skill (C++ lane, gating) and quality-lab / /audio-compare (advisory A/B with timestamped detectors).
  • Visual-compare tools in the CLAUDE.md tool registry, each with its Cannot see caveat.
  • trace-analysis skill when the question is "why is this slow" inside one process; a trace gives wall and CPU time per slice, which is a mechanism measure, unlike end-to-end wall time.

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

同じリポジトリのスキル

概要と使いどころ

aax

無料

Optional AAX support for Pulp, including developer-supplied Avid SDK setup, CMake enablement, DigiShell/AAX Validator workflows, and local AAX builds on macOS or Windows.

日本語の概要は準備中です。原文の説明を表示しています。

Generous-Corp/pulp222026年10月10日 更新

Configure, implement, and test Pulp's optional desktop Ableton Link tempo-sync adapter while preserving the developer-supplied SDK, licensing, realtime, latency-compensation, and no-install boundaries.

日本語の概要は準備中です。原文の説明を表示しています。

Generous-Corp/pulp222026年10月10日 更新

Maintain Pulp's installed design-time agent capability manifest and public-surface ledger. Use when adding, removing, renaming, or materially changing public audio, MIDI, signal, timebase, or sequence APIs; registering a new algorithm for generators; changing capability support or deprecation state; or repairing agent-capabilities freshness, schema, fingerprint, tombstone, or installed-SDK tests.

日本語の概要は準備中です。原文の説明を表示しています。

Generous-Corp/pulp222026年10月10日 更新

android

無料

Android platform development for Pulp — NDK cross-compilation, Oboe audio, Dawn/Skia GPU rendering, JNI bridge, touch interaction, emulator workflows, and end-to-end smoke validation. Covers build, deploy, debug, and the gotchas discovered during bringup.

日本語の概要は準備中です。原文の説明を表示しています。

Generous-Corp/pulp222026年10月10日 更新

ara

無料

Optional ARA support for Pulp, including developer-supplied ARA SDK setup, CMake enablement, adapter companion APIs, validation, and ARA-aware plugin implementation guidance.

日本語の概要は準備中です。原文の説明を表示しています。

Generous-Corp/pulp222026年10月10日 更新

The measurement surface for ALL Pulp DSP and audio-pipeline work — read it BEFORE writing or gating DSP, not only when something already sounds wrong. Covers the C++ harness (signal generators, metrics, assertions, RenderScenario, contracts), the offline Audio Doctor (magnitude/frequency response, THD/THD+N, phase/group delay), and their Python sibling the Audio Quality Lab (tools/audio/quality-lab — null residual + alignment, LTAS log-spectral distance, spectral flux/centroid, HNR, Theil-Sen drift slope, Kaiser-sinc resampling, license-guarded corpus, regression-net ratchet). TRIGGER on AUTHORING work — "build/design an oscillator/filter/synth/effect", "add a DSP module", "what should the acceptance gate be", "how do I measure aliasing / anti-aliasing / alias floor", "null against a reference", "is this DSP correct", "choose a tolerance", "golden/regression corpus for audio", "measure drift or jitter", "A/B two renders" — AND on DEBUGGING work — "is there sound / no audio / I hear nothing", "does this filter/compressor/synth/delay produce the right signal", "prove the DSP / prove the contract", "measure the frequency response", "what's the THD / is it distorting", "what's the group delay / phase response / measured latency", "magnitude response curve", "render a test tone and assert", "audio regression", "64-frame works but 128 is silent", "sample-rate change pitch-shifted it", "describe what's in this buffer", "audio doctor", "compare before/after a DSP refactor". Reach for this BEFORE hand-rolling any FFT, null test, alias measurement, pitch tracker, or golden-render script — most of it already exists in one of the two lanes. Test/tool layer over HeadlessHost — deterministic, no audio device, no speakers. Off the realtime thread entirely.

日本語の概要は準備中です。原文の説明を表示しています。

Generous-Corp/pulp222026年10月10日 更新

Generous-Corp のスキルをすべて見る

このスキルの問題を報告する