Look up any arxiv paper on alphaxiv.org to get a structured AI-generated overview. This is faster and more reliable than trying to read a raw PDF.
日本語の概要は準備中です。原文の説明を表示しています。
Developer reference for Senpai's retired direct-GitHub experiment logging workflow. Use only when auditing or migrating legacy research tracks.
インストールする前に、エージェントに与えられる指示の中身を確認できます。
This guide is not installed into live advisor or student runtimes. Current
students use the plugin's submit-experiment-results skill and typed result
tool instead.
The master agent creates an empty PR and hands it to you. Your job is to fill it out as you run trials and finalize it when you're done. The master reads these PRs to decide what to explore next — completeness and honest analysis matter more than polish.
<experiment-name> <best-wandb-run-id-if-there-is-a-positive-result>
experiment-name: the idea you explored (e.g. multi-scale-attn, surface-loss-reweight)best-wandb-run-id: W&B run ID of your best result (or most recent if all crashed)Fill this in as you go. Write the TL;DR last, once all trials are done.
# <agent-name>-<experiment-name>-<best-wandb-run-id>
## TL;DR - Result summary
<2-4 sentences: what you tried, what happened overall, and the bottom line verdict>
<!-- Fill in the block below only if a trial beat the baseline. Otherwise delete it. -->
**Main Hypothesis:** <what you believed would work and why>
**Key Result:** val_loss=X.XX | surf_Ux=X.XX | surf_p=XX.X (vs baseline: val_loss=X.XX)
**wandb:** <url to the best run>
## Trials run
### <YYYY-MM-DD_HH-MM> <wandb_run_id> <trial-name>
**Hypothesis:** <what you expected this specific run to show>
**Background research:** <relevant context — papers, similar approaches, why you thought this might work>
**Result:** val_loss=X.XX | surf_Ux=X.XX | surf_p=XX.X | memory=XX.XGB | status=keep/discard/crash
**wandb:** <wandb run url>
**Conclusion:** <what this result tells you — did the hypothesis hold? what would you try next?>
### <YYYY-MM-DD_HH-MM> <wandb_run_id> <trial-name>
...
Trials are ordered chronologically by start time, oldest first.
Starting the PR early lets the master see in-progress results and potentially redirect you before you've finished all planned trials.
まだレビューはありません。使ってみた感想をお寄せください。
概要と使いどころ
Look up any arxiv paper on alphaxiv.org to get a structured AI-generated overview. This is faster and more reliable than trying to read a raw PDF.
日本語の概要は準備中です。原文の説明を表示しています。
Operator-side analysis of historical ML experiment PRs in Senpai research tracks. Use this skill whenever the user asks to: analyze experiments, categorize PRs, bucket experiments, summarize what's been tried, understand experiment history, review merged vs closed results, or asks "what experiments have we run / worked / failed". Also triggers for: "pull the latest experiments", "what's been tried so far", "category breakdown of PRs", "which experiments succeeded", "noam track analysis". When a branch name is mentioned (e.g. "noam branch", "on the noam branch"), pass it as the base branch to scope the fetch to just those PRs.
日本語の概要は準備中です。原文の説明を表示しています。
Create a typed assignment branch and draft PR for one student. Use when the advisor has a concrete hypothesis and the student has no open `status:wip` or `status:review` assignment.
日本語の概要は準備中です。原文の説明を表示しています。
Create or improve a Senpai target repository's program.md. Use this skill whenever the user wants to point Senpai at a fresh ML or research target repository, define the research objective, primary metric, benchmark contract, allowed edit boundaries, W&B reporting contract, or prepare a repo for autonomous advisor/student experiment loops.
日本語の概要は準備中です。原文の説明を表示しています。
Open GitHub Issues for human input and respond to the researcher team. Use this skill whenever you need to handle a human_issue event, respond to human issues, ask humans a question, or check team communications. Also triggers for: "any human messages?", "check issues", "respond to humans".
日本語の概要は準備中です。原文の説明を表示しています。
Choose, configure, launch, and collect bounded subagents for research and engineering decisions. Root agents and delegation-capable subagents should read this before delegating: every task requires an explicit model tier, and high-leverage work such as research ideation, round planning, plateau pivots, large research reviews, hard optimization, disputed evidence, and expensive experiment portfolios requires frontier judgment.
日本語の概要は準備中です。原文の説明を表示しています。