Look up any arxiv paper on alphaxiv.org to get a structured AI-generated overview. This is faster and more reliable than trying to read a raw PDF.
日本語の概要は準備中です。原文の説明を表示しています。
Operator-side guide for generating a target-specific training curve chart for a legacy experiment PR. Use when auditing or presenting historical runs.
インストール方法を見るインストールする前に、エージェントに与えられる指示の中身を確認できます。
This guide is not installed into live advisor or student runtimes. Target repositories that require experiment charts should provide their own project skill with the correct metrics and plotting code.
Generate a comparison chart so a reviewer can see the training dynamics at a glance—not just the final numbers, but how the experiment got there. A bolded best-run line and a properly scaled y-axis make the story immediately readable, even if some runs diverged.
This skill takes about 30 seconds. It's worth it.
## Baseline, look for the W&B run: \xxxxxxxx`` line.WANDB_ENTITY and WANDB_PROJECT in the current environment.uv run .agents/skills/plot-experiment-charts/scripts/plot_training_curves.py \
--baseline <baseline-run-id> \
--runs <your-run-id-1>,<your-run-id-2>,...
The script automatically:
beta_scan_history (fast, batched)val_in_dist/mae_surf_p and bolds ittraining_curves.png in the current directoryIf you want to override which run gets bolded, pass --bold <run-id>.
Full flag reference:
| Flag | Default | Notes |
|---|---|---|
--baseline | required | 8-char W&B run ID of the current best baseline |
--runs | required | Comma-separated run IDs for your experiments (1–8) |
--bold | auto | Override which run gets the bold treatment |
--output | training_curves.png | Output filename |
--entity | $WANDB_ENTITY | W&B entity |
--project | $WANDB_PROJECT | W&B project |
git add train.py training_curves.png
git commit -m "<your experiment description>"
git push origin <branch>
The chart lives on the experiment branch. It's visible during review — which is the only time it matters. After the PR is squash-merged and the branch deleted, the image in the archived PR body will show as broken, but by then the advisor has already reviewed it.
The script prints a raw GitHub URL when it finishes. Copy it and add this to the ## Results section of the PR body:
## Results

| Metric | Baseline | This run |
| ... |
Put the chart before the metrics table so the advisor sees the curves first, then the numbers.
Two panels, side by side:
val_in_dist/mae_surf_p — surface pressure MAE on in-distribution data. This is the primary metric. Lower is better.val/loss — combined validation loss across all splits. Lower is better.The black dashed line is the baseline. Your runs are colored lines. The bold colored line is your best run — the one that would be a candidate for merging.
The y-axis is clamped at 3× the baseline's best value. If a run diverges far above that, it's clipped — that's fine, the advisor just needs to see "this diverged" rather than the exact value. The important region (near the baseline) stays readable.
The script skips runs with no history and prints a warning. Include the missing run IDs in your PR results section with a note about why they crashed — don't silently omit them.
"Run not found" — Double-check the run ID (8 alphanumeric chars, case-sensitive). The run ID is in the W&B run URL: wandb.ai/{entity}/{project}/runs/{run-id}.
"No data for key val_in_dist/mae_surf_p" — The run may have crashed before the first validation epoch. Check the run logs. Still include it in the chart call — the script handles empty runs gracefully.
Script not found — Run from the repo root: uv run .agents/skills/plot-experiment-charts/scripts/plot_training_curves.py .... If the helper script is not present in this repo checkout, generate the comparison chart manually from W&B history instead of blocking on the helper.
まだレビューはありません。使ってみた感想をお寄せください。
概要と使いどころ
Look up any arxiv paper on alphaxiv.org to get a structured AI-generated overview. This is faster and more reliable than trying to read a raw PDF.
日本語の概要は準備中です。原文の説明を表示しています。
Operator-side analysis of historical ML experiment PRs in Senpai research tracks. Use this skill whenever the user asks to: analyze experiments, categorize PRs, bucket experiments, summarize what's been tried, understand experiment history, review merged vs closed results, or asks "what experiments have we run / worked / failed". Also triggers for: "pull the latest experiments", "what's been tried so far", "category breakdown of PRs", "which experiments succeeded", "noam track analysis". When a branch name is mentioned (e.g. "noam branch", "on the noam branch"), pass it as the base branch to scope the fetch to just those PRs.
日本語の概要は準備中です。原文の説明を表示しています。
Create a typed assignment branch and draft PR for one student. Use when the advisor has a concrete hypothesis and the student has no open `status:wip` or `status:review` assignment.
日本語の概要は準備中です。原文の説明を表示しています。
Create or improve a Senpai target repository's program.md. Use this skill whenever the user wants to point Senpai at a fresh ML or research target repository, define the research objective, primary metric, benchmark contract, allowed edit boundaries, W&B reporting contract, or prepare a repo for autonomous advisor/student experiment loops.
日本語の概要は準備中です。原文の説明を表示しています。
Open GitHub Issues for human input and respond to the researcher team. Use this skill whenever you need to handle a human_issue event, respond to human issues, ask humans a question, or check team communications. Also triggers for: "any human messages?", "check issues", "respond to humans".
日本語の概要は準備中です。原文の説明を表示しています。
Choose, configure, launch, and collect bounded subagents for research and engineering decisions. Root agents and delegation-capable subagents should read this before delegating: every task requires an explicit model tier, and high-leverage work such as research ideation, round planning, plateau pivots, large research reviews, hard optimization, disputed evidence, and expensive experiment portfolios requires frontier judgment.
日本語の概要は準備中です。原文の説明を表示しています。