Use when adding support for a new model architecture to imp, porting a model family, or debugging a model that loads but produces wrong output - "add support for <model>", "new arch", loader detection, chat template, tokenizer parity, RoPE variant, "outputs garbage", "prompt-blind", "digits scrambled", "NaN logits", "describes a different picture", "does it fit in VRAM". Do NOT use for kernel performance (sm120-cuda-expert) or quant-format questions (quant-formats).
日本語の概要は準備中です。原文の説明を表示しています。
kekzl/imp☆ 442026年10月11日 更新
Use when benchmarking, profiling, or A/B-testing CUDA kernels or end-to-end perf in the imp inference engine on RTX 5090 (sm_120), including refreshing tests/perf_baseline.json or publishing numbers to docs/BENCHMARKS.md and the README. Triggers on "benchmark kernel", "profile cuda", "ncu", "nsys", "kernel timing", "kernel sum", "occupancy", "bandwidth bound", "compute bound", "roofline", "perf baseline", "is this regression real", "decode dropped", "aggregate throughput", "two-image A/B", "prefill kernel A/B". Do NOT use for writing/optimizing kernel code (sm120-cuda-expert) or output-quality checks (check-degeneration).
日本語の概要は準備中です。原文の説明を表示しています。
kekzl/imp☆ 442026年10月11日 更新
Use when building imp, running its test suite, checking CI status, or debugging build/test failures - "make build", "run the tests", "test-gpu", "verify-fast", GTEST_FILTER, Docker/CUDA toolchain, dependency bumps, "CI is red/blocked", "which gate failed", stale objects / segfault after a header edit, hook edits, docker-entrypoint env vars, determinism or perplexity checks. Do NOT use for benchmarking/profiling (benchmark-cuda) or output-quality batteries (check-degeneration).
日本語の概要は準備中です。原文の説明を表示しています。
kekzl/imp☆ 442026年10月11日 更新
Use when verifying that a model in the imp inference engine produces coherent output without repetition loops, token-stuck states, or state corruption across turns or streams. Triggers on "degenerates", "check degeneration", "repetition loop", "own own own", "stuck token", "empty content", "multi-turn regression", "does it still work", "NIAH", and after enabling CUDA graphs / changing forward pass / MoE routing / KV cache or KV dtype / GDN state or scan / sparse attention / PDL / speculation (MTP, n-gram) / ragged prefill / batched-decode kernels (smallm, producer quantize) / FA2 softmax.
日本語の概要は準備中です。原文の説明を表示しています。
kekzl/imp☆ 442026年10月11日 更新
Use when a question is about *structure* rather than text - who calls or launches a symbol, where it is defined, what a change would reach, whether something is dead, how a request gets from the API to a kernel. Triggers on "who calls", "who launches this kernel", "where is X defined", "what breaks if I change", "is this still used", "is this dead", "trace the path from X to Y", "blast radius", "what depends on this header". Do NOT use for free-text search (`rg` is better and cheaper) or to open a file whose path you already know.
日本語の概要は準備中です。原文の説明を表示しています。
kekzl/imp☆ 442026年10月11日 更新
Use when writing, moving or auditing any .md in imp - deciding which file a paragraph belongs in, adding a doc, fixing a stale claim, or when docs_lint.py fails in CI. Covers the four reader layers (L0 README / L1 operators / L2 kernel devs / L3 agents), the HTML-comment metadata header, [PROV:] provenance, the single-source-of-truth map, generated perf blocks, which numbers may appear in prose, plan-doc closure. Triggers on "which doc does this go in", "docs lint failed", "add a doc", "this claim is stale", "update the README", "PROV block", "layer", "STALE.md". Do NOT use for the CHANGELOG or a release body (shipping-prs), or for code-comment accuracy (codebase-audit).
日本語の概要は準備中です。原文の説明を表示しています。
kekzl/imp☆ 442026年10月11日 更新