本文へ移動
cccskills
無料GitHub で公開

plot-experiment-charts

Operator-side guide for generating a target-specific training curve chart for a legacy experiment PR. Use when auditing or presenting historical runs.

インストール方法を見る

含まれるファイル(1)

  • SKILL.md4.7 KB

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

Plot Experiment Charts

This guide is not installed into live advisor or student runtimes. Target repositories that require experiment charts should provide their own project skill with the correct metrics and plotting code.

Generate a comparison chart so a reviewer can see the training dynamics at a glance—not just the final numbers, but how the experiment got there. A bolded best-run line and a properly scaled y-axis make the story immediately readable, even if some runs diverged.

This skill takes about 30 seconds. It's worth it.

What you need

  • Baseline W&B run ID: in the PR body under ## Baseline, look for the W&B run: \xxxxxxxx`` line.
  • Your own run IDs: the 8-character W&B IDs of the runs you just completed. Find them in the W&B run URLs or in the training output (the run ID is printed at launch).
  • W&B credentials: WANDB_ENTITY and WANDB_PROJECT in the current environment.

Step 1 — Run the script

uv run .agents/skills/plot-experiment-charts/scripts/plot_training_curves.py \
  --baseline <baseline-run-id> \
  --runs <your-run-id-1>,<your-run-id-2>,...

The script automatically:

  • Downloads the full training history for each run using W&B's beta_scan_history (fast, batched)
  • Detects which of your runs achieved the lowest final val_in_dist/mae_surf_p and bolds it
  • Clamps the y-axis so the baseline curve stays readable even if some runs diverged
  • Saves training_curves.png in the current directory
  • Prints the GitHub raw URL to embed

If you want to override which run gets bolded, pass --bold <run-id>.

Full flag reference:

FlagDefaultNotes
--baselinerequired8-char W&B run ID of the current best baseline
--runsrequiredComma-separated run IDs for your experiments (1–8)
--boldautoOverride which run gets the bold treatment
--outputtraining_curves.pngOutput filename
--entity$WANDB_ENTITYW&B entity
--project$WANDB_PROJECTW&B project

Step 2 — Commit the chart alongside train.py

git add train.py training_curves.png
git commit -m "<your experiment description>"
git push origin <branch>

The chart lives on the experiment branch. It's visible during review — which is the only time it matters. After the PR is squash-merged and the branch deleted, the image in the archived PR body will show as broken, but by then the advisor has already reviewed it.

Step 3 — Embed in the PR body

The script prints a raw GitHub URL when it finishes. Copy it and add this to the ## Results section of the PR body:

## Results

![Training curves](https://raw.githubusercontent.com/owner/repo/branch/training_curves.png)

| Metric | Baseline | This run |
| ... |

Put the chart before the metrics table so the advisor sees the curves first, then the numbers.

Reading the chart

Two panels, side by side:

  • Left: val_in_dist/mae_surf_p — surface pressure MAE on in-distribution data. This is the primary metric. Lower is better.
  • Right: val/loss — combined validation loss across all splits. Lower is better.

The black dashed line is the baseline. Your runs are colored lines. The bold colored line is your best run — the one that would be a candidate for merging.

The y-axis is clamped at 3× the baseline's best value. If a run diverges far above that, it's clipped — that's fine, the advisor just needs to see "this diverged" rather than the exact value. The important region (near the baseline) stays readable.

If a run crashed or never logged metrics

The script skips runs with no history and prints a warning. Include the missing run IDs in your PR results section with a note about why they crashed — don't silently omit them.

Troubleshooting

"Run not found" — Double-check the run ID (8 alphanumeric chars, case-sensitive). The run ID is in the W&B run URL: wandb.ai/{entity}/{project}/runs/{run-id}.

"No data for key val_in_dist/mae_surf_p" — The run may have crashed before the first validation epoch. Check the run logs. Still include it in the chart call — the script handles empty runs gracefully.

Script not found — Run from the repo root: uv run .agents/skills/plot-experiment-charts/scripts/plot_training_curves.py .... If the helper script is not present in this repo checkout, generate the comparison chart manually from W&B history instead of blocking on the helper.

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

同じリポジトリのスキル

概要と使いどころ

Look up any arxiv paper on alphaxiv.org to get a structured AI-generated overview. This is faster and more reliable than trying to read a raw PDF.

日本語の概要は準備中です。原文の説明を表示しています。

wandb/senpai372026年10月6日 更新

Operator-side analysis of historical ML experiment PRs in Senpai research tracks. Use this skill whenever the user asks to: analyze experiments, categorize PRs, bucket experiments, summarize what's been tried, understand experiment history, review merged vs closed results, or asks "what experiments have we run / worked / failed". Also triggers for: "pull the latest experiments", "what's been tried so far", "category breakdown of PRs", "which experiments succeeded", "noam track analysis". When a branch name is mentioned (e.g. "noam branch", "on the noam branch"), pass it as the base branch to scope the fetch to just those PRs.

日本語の概要は準備中です。原文の説明を表示しています。

wandb/senpai372026年10月6日 更新

Create a typed assignment branch and draft PR for one student. Use when the advisor has a concrete hypothesis and the student has no open `status:wip` or `status:review` assignment.

日本語の概要は準備中です。原文の説明を表示しています。

wandb/senpai372026年10月6日 更新

Create or improve a Senpai target repository's program.md. Use this skill whenever the user wants to point Senpai at a fresh ML or research target repository, define the research objective, primary metric, benchmark contract, allowed edit boundaries, W&B reporting contract, or prepare a repo for autonomous advisor/student experiment loops.

日本語の概要は準備中です。原文の説明を表示しています。

wandb/senpai372026年10月6日 更新

Open GitHub Issues for human input and respond to the researcher team. Use this skill whenever you need to handle a human_issue event, respond to human issues, ask humans a question, or check team communications. Also triggers for: "any human messages?", "check issues", "respond to humans".

日本語の概要は準備中です。原文の説明を表示しています。

wandb/senpai372026年10月6日 更新

Choose, configure, launch, and collect bounded subagents for research and engineering decisions. Root agents and delegation-capable subagents should read this before delegating: every task requires an explicit model tier, and high-leverage work such as research ideation, round planning, plateau pivots, large research reviews, hard optimization, disputed evidence, and expensive experiment portfolios requires frontier judgment.

日本語の概要は準備中です。原文の説明を表示しています。

wandb/senpai372026年10月6日 更新

wandb のスキルをすべて見る

このスキルの問題を報告する