本文へ移動
cccskills
無料GitHub で公開

codebase-audit

Use when auditing the imp codebase for structural debt, dead code, god-objects/files, duplication, flag sprawl, or deciding whether a cleanup is worth shipping - "structure audit", "tech debt", "is this still used", "dead code", "refactor for clarity", "should we split this file", "remove this flag", "file size gate red", "alloc sites gate red". Do NOT use for build/test mechanics (building-and-testing), perf/kernel work (benchmark-cuda / sm120-cuda-expert), or output quality (check-degeneration).

インストール方法を見る

含まれるファイル(1)

  • SKILL.md8.3 KB

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

Codebase audit & structural debt - imp

Two hard rules

  1. A raw finding is a candidate. Fan-out sweeps over-flag: on 2026-06-09 findings C1/C2/C4 were all REFUTED on inspection and several "dead flags/files" were live. Verify each against the code before reporting; the value is in the verification.
  2. Read docs/audit/SETTLED.md BEFORE generating hypotheses. The 2026-07-29 audit generated first: 8 of 13 hypotheses were already-collapsed duplication (its section 15) and 5 facts in its own brief were stale (9 vs 16 architectures, C++20 vs C++23, src/graph/ vs src/exec/; section 19). A candidate that contradicts a ledger entry needs that entry's ANCHOR disproven. Closing a finding updates BOTH the report status line and the ledger (scripts/check-release.sh section 1d fails CI on a stale open entry; F-6, S-10 ("still open" two days after #1209) and section G each went stale for days in the 07-29 campaign).

Verification recipes

ClaimCheckTrap that fooled a past audit
flag/symbol X is deadrg across src/ tests/ tools/dump_tokens, force_bos, bench.generate looked dead in src/, read in tools/imp-cli/main.cpp
config field never readgrep the leaf for a value-read in a conditional, excluding decl/parse/copya local var with the field's name hid the read (fp8_auto_legacy)
kernel/function deadtrace the call graph, not <<< sites; code-graph has NO launch edges right now (ccg enrich broken) so rg the launcherq4k_imma CACHE was dead; mmq_q4k_imma_tile/_reorder live via q4k_imma_prefill -> mmq_q4k_imma_gemm; only ~30 LOC dead, not "~1200"
env var deadgrep docker-entrypoint.sh tooIMP_KV_FP8 read nowhere in src/, translated by the entrypoint into --kv-fp8
god-file, split itsplit on conflation or compile-time isolation, not sizegdn.cu, gemm.cu, json_schema.cpp are one domain each; the 7 attention_paged* share attention_paged_common.cuh. Precedent: engine_scheduler.cpp 2230 -> 1291 + engine_prefill.cpp + engine_decode_pipeline.cpp (#1782); gdn_scan_chunkpar.cu -> K1 + gdn_scan_chunkpar_pass.cu + .cuh at the 600-LOC kernel line; sampling family -> executor_sampling.cu (#1790); smallm -> executor_gemm_smallm.cpp (#1793)
rewrite into a registry/templateneed a concrete bugthe weight_map.cpp if (!matched) ladder is readable; a mis-mapped name = garbage output
imp.conf key inertread in src/ AND tools/? C-API bridge at engine init?server.prefix_cache was parsed, unread, bridged (#636)
a log line proves the code pathcompare the printed number against the code that uses ita CORRECT log line named a number the code did not use (#1746/#1705); the sparse startup line printed the 16-block arithmetic while ACTIVE printed the real one (#1819)
a number in a .cu is wrongpull the arithmetic into a pure function + CPU tests at both operating pointssrc/exec/sparse_attn_geometry.h, 7 tests, 4 fail on the original defect (#1819)

Workflow

  1. Ledger first: docs/audit/SETTLED.md; memory axis: root AUDIT.md; structural: docs/audit/DEBT_LEDGER_2026_08_21.md (stale [allow] reasons, the "every gate is advisory" prehistory of the blocking gates). Re-derive every brief fact (SETTLED section F lists the five that were wrong, each with a counter-check).
  2. Fan out (Explore agents, grep) AGAINST the ledger.
  3. Verify each with the recipes; demote refuted ones immediately.
  4. Write up: dated ledger in docs/audit/ (DEBT_LEDGER_2026_08_21.md, AUDIT_ARCH_2026_07_29.md; older structural_debt_* in docs/archive/), append refutations with anchors to SETTLED.md ("Refuted (do not re-chase)"); memory findings go to root AUDIT.md (CONFIRMED / REFUTED / OPEN). Counts are evidence, not estimates.
  5. Ship as small PRs off origin/main, make build + make test-unit + hooks (skill building-and-testing, shipping-prs); behavior-sensitive removals also run check-degeneration.

Gates that shape audit findings (scripts/ci_static_gates.sh; all block inside Build since #1527; hooks run them in ~2 s)

GateToolRule
File sizetools/check_filesize.py, tools/filesize_thresholds.tomlcode LOC (comments/blanks stripped): kernel .cu warn 500 / hard 600, TU 600 / 800, header 500 / 700. [allow] entries { code_loc = N, reason = "..." } are a TWO-WAY ceiling since #1526 (drift either way fails; --update re-pins). Cost metric = recompile blast radius (one .cu = one ptxas TU), so a #included .cu is charged to its includer since #1905 — executor_attention.cu read 542 as a file and 1279 as the TU. Rationale per file: docs/audit/AUDIT_FILESIZE.md. The filesize group also runs check_determinism_sites.py, check_dead_inline_accessors.py, check_log_fatal.py
Function sizetools/check_function_size.py, tools/function_size_thresholds.tomlthe largest body in a file, warn 200 / hard 500 code LOC (p99 of 5830 functions is 199). Same two-way [allow] ceiling, keyed path::signature. Exists because a file's cohesion reason was covering an 884-line function inside an allowlisted "(c) one concern" file (#1905)
Test lanestools/check_test_lanes.py --reportown check since #1770; fails on a test in no ctest entry or a unit-lane test demoted vs the base; no-lane count printed with its delta
Alloc sitestools/check_alloc_sites.py + tools/alloc_allowlist.txt, tools/check_alloc_pairs.pyinvariant I1 (one module talks to the driver); fails on a new site AND a stale entry (#1479 left gdn.cu's entry behind). Counts SOURCE sites: the T2 slot pool removed ~96% of runtime allocations with a flat site count (AUDIT.md B34); runtime number via make check-alloc-interpose (steady_state_allocations(), never benchmark that build, AUDIT.md G16)
Dead inline accessorstools/check_dead_inline_accessors.py (make check-dead-inline), tools/dead_inline_allowlist.txtheader-inline definitions with no caller
Launch guardstools/check_launch_guards.pypost-launch check present
Kernel resourcesmake kernel-resources, tools/kernel_resource_baseline.txtregister/local-frame ratchet (#1549)
Doc citationsscripts/check_doc_citations.py .a file split moves line numbers and kills file:line in living docs (#1782 -> #1783); docs/audit/ itself is exempt

Priors: settled, do not re-flag

Canonical: docs/audit/SETTLED.md (collapsed duplication S-1..S-11, deliberate specialisations, hunted-and-absent negatives, findings that were themselves wrong). Not yet in the ledger:

  • Two config systems (RuntimeConfig vs ModelConfig::Overrides) are deliberate; keys use the typed binder ladder B/I/F/S(...) (#626); the config surface lives in src/core/config/*.h.
  • Structural-debt chain D1-D4 / C1-C8 archived in docs/archive/README.md, verdicts in docs/archive/housekeeping_2026_06_13.md. Audit #5 (docs/archive/structural_debt_2026_07_07.md, companion docs/archive/vram_audit_2026_07_07.md): server layer is the debt source, issues #888-#897; read its NOT-flagged list first. src/memory/vram_owned.h EXISTS since #1530 (once a hallucinated finding).
  • Memory (AUDIT.md, 2026-07-29): "engine teardowns leak ~15 GiB" is REFUTED as a leak: every CUDA release works, WSL2/WDDM never returns a process's peak commitment. Traps: cudaMalloc success proves nothing (28 GiB at 22.6 GiB free; ~1530 vs ~237 GB/s tells), cudaMemPoolAttrReservedMemCurrent drops to 0 on trim, assert on UsedMemCurrent.
  • Serving-loop classes closed by measurement: host turnaround (#1791), elementwise/residual fusion (#1793), prefill||decode overlap (#1792), H2D consolidation (#1834). Roadmap item 0(c) complete.
  • A refactor that touches exec/executor.h needs a forced rebuild (grep -rl exec/executor.h src tools tests | xargs touch && make dev): stale objects segfault and look like the refactor.

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

同じリポジトリのスキル

概要と使いどころ

Use when adding support for a new model architecture to imp, porting a model family, or debugging a model that loads but produces wrong output - "add support for <model>", "new arch", loader detection, chat template, tokenizer parity, RoPE variant, "outputs garbage", "prompt-blind", "digits scrambled", "NaN logits", "describes a different picture", "does it fit in VRAM". Do NOT use for kernel performance (sm120-cuda-expert) or quant-format questions (quant-formats).

日本語の概要は準備中です。原文の説明を表示しています。

kekzl/imp442026年10月11日 更新

Use when benchmarking, profiling, or A/B-testing CUDA kernels or end-to-end perf in the imp inference engine on RTX 5090 (sm_120), including refreshing tests/perf_baseline.json or publishing numbers to docs/BENCHMARKS.md and the README. Triggers on "benchmark kernel", "profile cuda", "ncu", "nsys", "kernel timing", "kernel sum", "occupancy", "bandwidth bound", "compute bound", "roofline", "perf baseline", "is this regression real", "decode dropped", "aggregate throughput", "two-image A/B", "prefill kernel A/B". Do NOT use for writing/optimizing kernel code (sm120-cuda-expert) or output-quality checks (check-degeneration).

日本語の概要は準備中です。原文の説明を表示しています。

kekzl/imp442026年10月11日 更新

Use when building imp, running its test suite, checking CI status, or debugging build/test failures - "make build", "run the tests", "test-gpu", "verify-fast", GTEST_FILTER, Docker/CUDA toolchain, dependency bumps, "CI is red/blocked", "which gate failed", stale objects / segfault after a header edit, hook edits, docker-entrypoint env vars, determinism or perplexity checks. Do NOT use for benchmarking/profiling (benchmark-cuda) or output-quality batteries (check-degeneration).

日本語の概要は準備中です。原文の説明を表示しています。

kekzl/imp442026年10月11日 更新

Use when verifying that a model in the imp inference engine produces coherent output without repetition loops, token-stuck states, or state corruption across turns or streams. Triggers on "degenerates", "check degeneration", "repetition loop", "own own own", "stuck token", "empty content", "multi-turn regression", "does it still work", "NIAH", and after enabling CUDA graphs / changing forward pass / MoE routing / KV cache or KV dtype / GDN state or scan / sparse attention / PDL / speculation (MTP, n-gram) / ragged prefill / batched-decode kernels (smallm, producer quantize) / FA2 softmax.

日本語の概要は準備中です。原文の説明を表示しています。

kekzl/imp442026年10月11日 更新

Use when a question is about *structure* rather than text - who calls or launches a symbol, where it is defined, what a change would reach, whether something is dead, how a request gets from the API to a kernel. Triggers on "who calls", "who launches this kernel", "where is X defined", "what breaks if I change", "is this still used", "is this dead", "trace the path from X to Y", "blast radius", "what depends on this header". Do NOT use for free-text search (`rg` is better and cheaper) or to open a file whose path you already know.

日本語の概要は準備中です。原文の説明を表示しています。

kekzl/imp442026年10月11日 更新

Use when writing, moving or auditing any .md in imp - deciding which file a paragraph belongs in, adding a doc, fixing a stale claim, or when docs_lint.py fails in CI. Covers the four reader layers (L0 README / L1 operators / L2 kernel devs / L3 agents), the HTML-comment metadata header, [PROV:] provenance, the single-source-of-truth map, generated perf blocks, which numbers may appear in prose, plan-doc closure. Triggers on "which doc does this go in", "docs lint failed", "add a doc", "this claim is stale", "update the README", "PROV block", "layer", "STALE.md". Do NOT use for the CHANGELOG or a release body (shipping-prs), or for code-comment accuracy (codebase-audit).

日本語の概要は準備中です。原文の説明を表示しています。

kekzl/imp442026年10月11日 更新

kekzl のスキルをすべて見る

このスキルの問題を報告する