本文へ移動
cccskills
無料GitHub で公開

autonomous-gpu-kernel-timeline

Let AKA autonomously add, run, inspect, and revise intra-kernel timeline probes for standalone CUDA/inline PTX or CuTe DSL when ordinary benchmark, NSYS, or NCU evidence cannot answer a specific kernel-internal timing question.

インストール方法を見る

含まれるファイル(8)

  • SKILL.md4.9 KB
  • backends/cuda_backend/adapter.py3.7 KB
  • backends/cuda_backend/atrex_timeline.cuh7.8 KB
  • backends/cuda_backend/test_backend.cu7.7 KB
  • backends/cutedsl_backend/adapter.py16.3 KB
  • references/cuda-backend.md4.1 KB
  • references/iket-quickstart.md5.5 KB
  • scripts/timeline.py67.9 KB

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

Autonomous GPU Kernel Timeline

Use this inside an AKA optimization episode only after a correct runnable kernel and representative workload exist. Do not trigger it when aggregate kernel timing or ordinary profiler evidence already answers the question.

Ownership

  • AKA chooses the hypothesis, sites, writer roles, density, boundary semantics, and number of retries.
  • The backend records and decodes those choices without silent identity, capacity, or pairing errors.
  • The final validator rejects known factual failures; it does not decide whether AKA chose the best algorithm boundary or interpretation.

Route

Autonomous loop

  1. State one falsifiable timing question and begin with the fewest useful sites.
  2. Save the current clean kernel.py content, then create an instrumented working snapshot in the same episode worktree. This snapshot is not a handoff candidate and must not enter promotion or stall accounting.
  3. Run representative correctness through the immutable evaluator and capture through the campaign sandbox. For final evidence, pass the sandbox-owned .atrex_long_horizon/evaluations.jsonl as --correctness-evidence; a model-supplied --correctness passed is only an exploration note and cannot produce decision_grade. Use scripts/timeline.py to validate and export the returned evidence. If the remote command reads backend files, pass --input skills/autonomous-gpu-kernel-timeline/backends/<backend> to tools/sandbox.py; sync only the attempt-specific directory.
  4. Read the summary and Perfetto trace. Remove, move, or refine probes and repeat when the evidence is insufficient, semantically misplaced, or too intrusive. Name every reported interval as start_site -> end_site and recompute its numbers from canonical events rather than trusting a prose label. A local timeline delta is mechanism evidence, not a substitute for probe-free end-to-end ABBA. Reviewer feedback is optional.
  5. Before replacing the snapshot, preserve its source or reversible patch and its content hash.
  6. Restore a probe-free kernel, implement the optimization, and use the normal evaluator for final correctness and performance. Re-instrument the new clean state only when confirmation is useful.

For a final perturbation claim, use scripts/timeline.py measure inside one GPU allocation. Give it baseline/instrumented commands as JSON argv arrays, the exact sources, and materialized binaries when the compiler exposes them. JIT-only CuTe DSL measurements need not invent a binary artifact. Each command must perform the requested warmup and iterations, synchronize the device, check the full representative output, and emit exactly one line of this form:

__ATREX_TIMELINE_SAMPLE__={"latency_ms":1.0,"correctness":"passed","synchronized":true,"workload_identity":"...","device_identity":{"uuid":"..."},"warmup":10,"iterations":100}

The helper runs each sample in a fresh process, defaults to ABBA followed by BAAB, rejects workload or device drift, and writes the raw schedule and samples. Pass that artifact to capture/export with --measurement; validate rechecks its hashes and recomputes the medians and overhead. The sample's correctness field rejects a bad timing run but does not replace the immutable evaluator record required for final evidence.

Hard boundaries

  • Never modify profile_driver.py, evaluators, ground truth, or other protected paths.
  • Never hand off or promote an instrumented snapshot. Only a probe-free kernel may become the episode candidate_commit == HEAD.
  • Do not combine events from different launches into one apparent execution.
  • Construct exactly one recorder per selected owner per launch and reuse it for that owner's events; duplicate ownership is a capture failure, not a sampling policy.
  • Reject overflow, truncation, invalid owner/site identity, deterministic range mismatch, stale source/binary/workload provenance, and unrecomputable numeric claims.
  • Never classify final evidence as decision_grade without a matching sandbox evaluator record.
  • The default low-perturbation target is 10%. A higher-overhead trace may remain diagnostic when it is labelled honestly; the threshold is hard only when the task explicitly requires it.
  • Do not add fixed site catalogs, fixed coverage quotas, mandatory reviewer approval, or a separate orchestration state machine.

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

同じリポジトリのスキル

概要と使いどころ

gen-plan

無料

Generate a structured implementation plan from an evidence draft. Validate paths, obtain configured independent Codex and Qoder reviews, synthesize available advice against repository evidence, preserve the draft, and produce testable acceptance criteria and validation steps.

日本語の概要は準備中です。原文の説明を表示しています。

alibaba/atrex-kernel-agent1672026年10月10日 更新

Learn the target framework from enabled knowledge tools and implement a baseline GPU kernel. Use this skill to understand compute semantics, determine the target platform and framework, search reference implementations, and produce a correct V0 baseline with performance records for later profile-driven optimization.

日本語の概要は準備中です。原文の説明を表示しています。

alibaba/atrex-kernel-agent1672026年10月10日 更新

Run the evidence loop of one long-horizon GPU kernel optimization episode. Use this skill to reconstruct the incumbent, profile and localize a bottleneck, research progressively, plan one coherent direction, implement and repair, validate development correctness and performance, and record every decisive experiment in the episode journal.

日本語の概要は準備中です。原文の説明を表示しています。

alibaba/atrex-kernel-agent1672026年10月10日 更新

Mine a per-kernel optimization trace — a git repository capturing successive versions of one kernel being optimized — into structured, gate-validated optimization-experience records for the GPU kernel wiki. Use when asked to distill an optimization run, a kernel_opt trace directory or a version ladder into wiki records; to report what such a run actually achieved; to build, extend, re-run or validate the staging store behind those records; or to explain how a trace-derived record's number, snippet or provenance was established.

日本語の概要は準備中です。原文の説明を表示しています。

alibaba/atrex-kernel-agent1672026年10月10日 更新

Choose and run ACU-only, adaptive PPU in-kernel timeline, or optional bounded joint analysis for a PPU kernel. Use for device-level bottleneck diagnosis, kernel-internal critical-path questions, or evidence that genuinely needs both; do not require all three modes.

日本語の概要は準備中です。原文の説明を表示しています。

alibaba/atrex-kernel-agent1672026年10月10日 更新

Quickly extract a typed GPU-Wiki query intent from prose. Never inspect or query a knowledge store.

日本語の概要は準備中です。原文の説明を表示しています。

alibaba/atrex-kernel-agent1672026年10月10日 更新

alibaba のスキルをすべて見る

このスキルの問題を報告する