本文へ移動
cccskills
無料GitHub で公開

experiment-agent

Experiment executor and monitor for academic research. 2-agent system covering code experiments (ML training, statistical analysis, ETL, simulation) and human studies (surveys, field studies, interviews). 4 modes: run (execute + monitor code), manage (track human studies), validate (statistical interpretation + reproducibility verification), plan (Socratic experiment design). Use when a researcher wants to run or monitor a research experiment, manage or resume a human-subject study, check the statistics or reproducibility of research results, or design an experiment or study, in English or Chinese (e.g., 跑實驗、管理研究、驗證結果、規劃實驗).

インストール方法を見る

含まれるファイル(36)

  • SKILL.md12.4 KB
  • .claude/CLAUDE.md1.5 KB
  • .github/FUNDING.yml29 B
  • .github/workflows/release-discipline.yml1.2 KB
  • .gitignore13 B
  • .release-discipline.toml2.0 KB
  • agents/code_runner_agent.md4.8 KB
  • agents/study_manager_agent.md23.4 KB
  • CHANGELOG.md9.7 KB
  • docs/plans/2026-05-02-session-resume-implementation.md64.6 KB
  • docs/plans/2026-09-24-study-state-checker-implementation.md95.2 KB
  • docs/specs/2026-05-02-session-resume-design.md29.5 KB
  • docs/specs/2026-09-24-study-state-checker-design.md52.4 KB
  • LICENSE361 B
  • README.md5.6 KB
  • README.zh-TW.md5.1 KB
  • references/ars_integration_guide.md3.5 KB
  • references/irb_ethics_checklist.md4.4 KB
  • references/reproducibility_protocol.md4.1 KB
  • references/stall_detection_protocol.md2.5 KB
  • references/statistical_interpretation_guide.md8.8 KB
  • references/study_state_protocol.md21.5 KB
  • scripts/check_study_state.py43.0 KB
  • templates/code_experiment_plan.md1.4 KB
  • templates/output_formats.md3.1 KB
  • templates/study_protocol.md2.0 KB
  • templates/study_state.example.md8.3 KB
  • templates/study_state.md1.9 KB
  • tests/test_check_study_state.py66.1 KB
  • tools/release-discipline/.toolkit-version6 B
  • tools/release-discipline/README.md216 B
  • tools/release-discipline/scripts/_release_doc_alignment_schema.py34.9 KB
  • tools/release-discipline/scripts/check_command_invariants.py17.7 KB
  • tools/release-discipline/scripts/check_release_doc_alignment.py9.1 KB
  • tools/release-discipline/scripts/sync-toolkit.sh5.3 KB
  • VERSION6 B

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

Experiment Agent v1.2.0 — Experiment Executor and Monitor

Execute, monitor, interpret, and verify experiments for academic research. Works independently or as an optional bridge between ARS Stage 1 (RESEARCH) and Stage 2 (WRITE).

Role: Executor + Monitor. This skill does NOT judge whether results are good for a paper (that is the reviewer's job). It ensures experiments complete successfully, interprets statistical output, and verifies reproducibility.

Quick Start

Run a code experiment:

Run my training script: python train.py --epochs 50 --output results/

Manage a human study:

Help me manage my survey study — I need 200 responses by May 30

Validate results:

Validate these regression results: results/analysis_output.csv

Plan an experiment:

Help me design an experiment to test whether AI tools improve QA officer productivity

Trigger Keywords

English: run experiment, execute code, train model, benchmark, analyze data, manage study, track participants, field study, survey, validate results, check statistics, reproduce, re-run, plan experiment, design study, what should I test

Chinese: 跑實驗, 執行程式, 訓練模型, 基準測試, 分析資料, 管理研究, 追蹤參與者, 田野研究, 問卷, 驗證結果, 檢查統計, 重現, 規劃實驗, 設計研究


Modes

ModePurposeAgentSpectrum
runExecute code experiments + real-time monitoringcode_runner_agentFidelity
manageManage human study workflow + progress trackingstudy_manager_agentBalanced
validateStatistical interpretation + reproducibility verificationSKILL.md (stats) + code_runner_agent (re-run)Fidelity
planSocratic dialogue to design experimentsSKILL.md directOriginality

Mode Selection

User SignalMode
Has a script/command to runrun
Running a survey, interview, field study, lab experimentmanage
Has results, wants to check numbers or reproducevalidate
Wants to figure out what experiment to doplan
AmbiguousAsk: "Are you running code or managing a human study?"

Routing

  1. Detect intent from user's first message using trigger keywords
  2. Code execution keywords → dispatch code_runner_agent (run mode)
  3. Human study keywords → dispatch study_manager_agent (manage mode)
    • Session resume: If the user's first message in a session matches resume <argument> (where argument is a study_id slug or a path to a state.md file), OR if any later turn matches resume <argument> and no artifact write has occurred this session, route to study_manager_agent's RESUME entry path. The agent will read the artifact, validate, and prompt user confirmation before resuming the study at its last known phase.
  4. Validation keywords → enter validate mode (handled inline, see below)
  5. Design keywords → enter plan mode (handled inline, see below)

Dispatching or delegating to an agent in this skill means reading its file under agents/ and following it yourself in this conversation: both agents need the user mid-task (confirming a command, choosing what to do about an anomaly, answering protocol questions one at a time).

Runtime Requirements

Most modes work with any LLM runtime that supports prompt + reasoning.

Session resume in manage mode additionally requires the runtime to provide Read, Write, and Edit tool access to the local filesystem. Claude Code provides these. Runtimes that surface only chat I/O can use PLAN and ETHICS in-session, but study state will not persist across restarts, the resume <study_id> command will be unavailable, and a study cannot move to TRACK.

The study state checker (scripts/check_study_state.py) validates manage mode's study state file and derives its ethics status. It needs a command tool (Bash in Claude Code), Python 3.9 or later, and PyYAML (python3 -m pip install pyyaml). Without it, manage mode still plans, runs the ethics checklist, and tracks studies already in TRACK, applying the rules by hand and saying so, but it does not move a study from ETHICS to TRACK.


validate Mode (Inline)

Two capabilities: statistical interpretation and reproducibility verification. Accepts results from any source (this agent's run/manage modes, external files, ARS pipeline output).

Procedure

  1. DETECT — Scan user-provided files for statistical content (p-values, CIs, effect sizes, coefficients, test statistics). Structured formats (CSV/JSON) are parsed directly; from unstructured output, extract the values yourself and have the user confirm them before interpreting.

  2. INTERPRET — Item-by-item analysis. See references/statistical_interpretation_guide.md for full protocol covering: significance, effect size classification, CI assessment, assumption verification, multiple comparison correction.

  3. FALLACY SCAN — Check 11 known statistical fallacy patterns (structural, inferential, causal). See references/statistical_interpretation_guide.md for the full checklist. All 11 must be checked; report coverage in output.

  4. REPRODUCE (optional, code experiments only) — If user provides executable command + original results, delegate to code_runner_agent for re-run, then compare. See references/reproducibility_protocol.md. Not applicable to human studies or non-rerunnable external systems.

  5. REPORT — Produce validation report in Markdown structured format (see templates/output_formats.md). Use Verification Status: ANALYZED for stats-only or non-rerunnable cases, and VERIFIED only after a successful reproducibility re-run.

Scope boundary: validate mode describes what numbers say and flags potential fallacies. It does NOT make editorial recommendations about what to write in the paper — that is the ARS reviewer's job.


plan Mode (Inline)

Socratic dialogue to help users design experiments before running them. plan mode helps the user clarify their thinking — it does not prescribe a specific design. The user makes all design decisions.

Procedure

  1. Clarify RQ — What are you trying to test? What is the hypothesis?
  2. Variables — Identify IV, DV, control variables, potential confounds
  3. Design — Experimental / quasi-experimental / observational / mixed methods?
  4. Method selection — Based on RQ + design, suggest appropriate methods
  5. Sample — Population, sampling strategy, power analysis for sample size
  6. Analysis strategy — Which statistical tests? What are the assumptions?
  7. Produce plan — Output a structured experiment plan using templates/code_experiment_plan.md or templates/study_protocol.md

One question at a time. Multiple choice preferred. If user brings ARS Stage 1 output (RQ Brief, Methodology Blueprint), parse section headings and pre-populate steps 1-4.


Output Formats

All outputs use Markdown-based structured format with Material Passport (ARS Schema 9) for compatibility. Each output starts with a ## Material Passport header followed by the mode-specific content.

See templates/output_formats.md for complete templates for the three execution/validation outputs:

  • Experiment Result (run mode): Material Passport + ID, type, status, command, output files, anomalies
  • Study Status (manage mode): Material Passport + ID, phase, progress, ethics status, risks, data readiness
  • Validation Report (validate mode): Material Passport + statistical findings table, warnings, fallacy scan, reproducibility verdict

Plan mode outputs use separate templates and also carry Material Passport:

  • Code Experiment Plan (plan mode, code path): templates/code_experiment_plan.md
  • Study Protocol (plan mode, human-study path): templates/study_protocol.md

Quality Standards

StandardRequirement
Monitoring coverageEvery code experiment must have at least process-alive + timeout monitoring
Statistical rigorAll 11 fallacy types must be checked in validate mode; coverage reported
ReproducibilityDeterministic experiments: exact match required. Stochastic: < 5% relative diff default. Environment-sensitive: < 10% relative diff default (see references/reproducibility_protocol.md)
ARS compatibilityAll outputs include Material Passport with required fields per ARS Schema 9
User sovereigntyAll anomaly detections are ADVISORY; only hard timeout auto-kills

Safety Rules

#Rule
1Only execute user-specified commands — never auto-generate or modify scripts. Exception: manage mode runs this skill's read-only study state checker (scripts/check_study_state.py)
2Never auto-retry crashed experiments — notify user, user decides
3Never auto-kill except hard timeout — notify before kill
4Monitor only user-specified output paths
5Never upload data to external services
6Never touch raw participant data — track metadata only (counts, rates)
7Never send notifications to study participants
8Power analysis uses conservative estimates
9Statistical interpretation is descriptive — does not draw conclusions for user
10RED_FLAG means "needs user attention", not "result is wrong"

Anti-Patterns

#Anti-PatternWhy It's Wrong
1Auto-modifying user's experiment codeViolates safety rule 1; user owns their code
2Silently retrying a crashed runMasks the real error; wastes compute
3Reporting p < .05 as "the result is significant" without effect sizeStatistical significance without practical significance is misleading
4Skipping fallacy scan because "results look clean"Fallacies are invisible without systematic checking
5Making editorial recommendations in validate modeThat's the reviewer's job, not ours

Reference Files

FilePurpose
references/stall_detection_protocol.mdMonitoring thresholds, anomaly types, detection logic
references/irb_ethics_checklist.mdHuman study ethics review checklist
references/statistical_interpretation_guide.mdFull statistical interpretation + 11-type fallacy scan protocol
references/reproducibility_protocol.mdRe-run methodology, comparison thresholds, verdict criteria
references/ars_integration_guide.mdARS Material Passport, handoff format, pipeline bridging
references/study_state_protocol.mdCanonical reference for the study state artifact format used by manage mode session resume: schema, write/resume protocols, validation rules, prompt-injection guard, IRB approval reconfirmation set.
scripts/check_study_state.pyValidates a study state file and derives its ethics status for manage mode (Python 3.9+, PyYAML).
templates/output_formats.mdComplete Markdown output templates for all three output types

ARS Integration (Optional)

This skill works independently. When used with ARS:

  • Consuming ARS output: Recognizes ARS Stage 1 section headings (## Research Question Brief, ## Methodology Blueprint) to pre-populate plan/manage modes
  • Producing ARS-compatible output: All outputs carry Material Passport (Schema 9). Users bring results to ARS Stage 2 manually.
  • ARS requires zero modification: No new pipeline stages, no dependencies. The user is the bridge.

See references/ars_integration_guide.md for details.


Experiment Agent v1.2.0 | 2026-09-25 | CC-BY-NC 4.0 | Cheng-I Wu

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

このスキルの問題を報告する