本文へ移動
cccskills
無料GitHub で公開

experiment

Rules as hypotheses: falsifiable tests, confidence updates, graduate or kill. Triggers: experiment, hypothesis, prove, test rule, validate methodology, evidence.

インストール方法を見る

含まれるファイル(3)

  • SKILL.md9.8 KB
  • agents/openai.yaml217 B
  • reference/experiment-research.md6.2 KB

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

<skill id="experiment"> <purpose> Rules without evidence are superstitions. One invocation, no subcommands: the engine figures out what the hypothesis system needs — seeds, tests, graduates, or kills.
  • No hypotheses? Seeds them from CLAUDE.md.
  • Hypotheses exist? Picks the most uncertain, designs an experiment, runs it.
  • Evidence accumulating? Graduates proven rules, kills disproven ones.
  • Everything tested? Reports and stops.

Rules that survive become convictions. Rules that fail become learnings. </purpose>

<reference> Deep patterns, SQL schema, domain design templates, confidence-scoring examples: skills/experiment/reference/experiment-research.md </reference>

<skill_load> always: skills/quality/SKILL.md, skills/build/reference/testing.md </skill_load>

<on_start>

agentdb read-start
agentdb emit command "experiment-start" "" '{}'

</on_start>

<cycle id="experiment" max_iterations="20"> <phase id="sense" name="SENSE — Read the State"> Autonomous entry point. Determine what the system needs.
```bash
agentdb hypothesis list 2>/dev/null
```

Decision tree (no human input needed):
1. No hypotheses table or empty? → go to SEED phase.
2. Hypotheses exist but all are unproven? → go to PICK phase.
3. Mix of tested/untested? → go to PICK phase (prioritize untested).
4. All have >= 3 experiments? → go to JUDGE phase.
5. Graduation/kill candidates exist? → go to EVOLVE phase.
</phase> <phase id="seed" name="SEED — Extract Rules as Hypotheses" trigger="sense.no_hypotheses"> Scan every CLAUDE.md in the project hierarchy + rules/*.md + skills/*/SKILL.md.
Parse rule-like patterns:
- Imperative: "Always X", "Never Y", "Prefer Z", "Must W"
- Anti-patterns: block actions, "Don't", "Forbidden"
- Assertions: "X before Y", "X is better than Y"
- Quantitative claims: "reduces by X%", "takes N minutes"
- Conditional: "If X then Y", "When X, do Y"

For each rule:
```bash
agentdb hypothesis add "<statement>" --domain <auto-classify> --source "<file:line>"
```

Domain auto-classification by keyword:
- research, anti-pattern, prior work → methodology
- parallel, agent, spawn, tier → coordination
- test, coverage, edge case, mock → testing
- commit, branch, merge, PR → git
- secret, validation, auth, injection → security
- measure, optimize, latency, profile → performance
- Big 5, review, quality → quality
- module, interface, coupling → architecture

Deduplicate: skip if near-identical statement already exists.
Log count, then immediately proceed to PICK. No pause.

```bash
agentdb emit command "experiment-seed" "" '{"seeded":N}'
```
</phase> <phase id="pick" name="PICK — Select Next Hypothesis"> Choose the hypothesis that will produce the most information.
Priority order:
1. **Most uncertain**: confidence closest to 0.5 (maximum ignorance — any experiment is maximally informative)
2. **Least tested**: fewest total experiments (break ties)
3. **Highest impact domain**: methodology > coordination > security > testing > quality > git > architecture > performance

```sql
SELECT id, statement, domain, confidence, evidence_for + evidence_against as total_evidence
FROM hypotheses
WHERE status NOT IN ('graduated', 'refuted')
ORDER BY ABS(confidence - 0.5) ASC, total_evidence ASC
LIMIT 1;
```
</phase> <phase id="design" name="DESIGN — Create the Experiment"> Autonomously design the minimum viable experiment.
**Falsifiability gate: if no possible outcome could refute the hypothesis, redesign.**
Every experiment defines BEFORE running: method, quantitative measurement, control
condition (what happens WITHOUT the rule), pass_criteria, fail_criteria.

Choose the LIGHTEST experiment type that produces signal:

1. **HISTORICAL** (cheapest — query existing data):
   Query agentdb learnings, session outcomes, error patterns for evidence.
   Use when: agentdb has >= 10 sessions or >= 20 learnings in the domain.
2. **COMPARATIVE** (medium — run a real task two ways):
   Execute WITH the rule applied, then WITHOUT (or find prior without-cases).
   Measure: time, error count, rework, quality.
3. **ABLATION** (medium — remove the rule, observe):
   Temporarily ignore the rule during a real task. Record what breaks.
4. **OBSERVATIONAL** (passive — tag next N tasks):
   Flag the hypothesis; future relevant tasks collect evidence passively.
   Use when: active experimentation would be disruptive.

Minimum sample sizes: methodology/coordination/git/quality >= 3 comparisons;
testing >= 5 tasks per condition; security >= 50 fuzz inputs.

```bash
agentdb experiment add <H_ID> "<method>" "<measurement>" --pass-criteria "<criteria>"
```
</phase> <phase id="run" name="RUN — Execute and Observe"> Run the designed experiment. Record everything.
**Gate: the control condition was actually tested, not just assumed.**

- HISTORICAL: query agentdb with specific SQL; evidence = query result + interpretation.
- COMPARATIVE: execute the task (spawn agents if needed); evidence = measured delta.
- ABLATION: execute with the rule explicitly ignored; evidence = observed difference.
- OBSERVATIONAL: record the flag; skip to next hypothesis (no blocking).
</phase> <phase id="conclude" name="CONCLUDE — Verdict and Confidence Update"> Compare observations against pass/fail criteria. Issue verdict honestly: **supports** | **refutes** | **inconclusive**.
```bash
agentdb experiment verdict <EXP_ID> <supports|refutes|inconclusive> "<evidence summary>"
```

Confidence update (Bayesian, applied automatically by CLI):
- supports:     confidence += (1 - confidence) * 0.25
- refutes:      confidence -= confidence * 0.3
- inconclusive: no change

Evidence strings must be specific and measurable, never narrative.

Lifecycle transitions:
- unproven → testing: first experiment registered
- testing → supported: confidence >= 0.8 AND evidence_for >= 3 AND ratio >= 3:1
- testing → refuted: confidence < 0.2 AND evidence_against >= 2
- supported → graduated: human approval after sustained confidence
- refuted → killed: human approval to remove from rules
- any → unproven: rule is modified (resets all evidence)

```bash
agentdb learn pattern|failure "<what we learned>" "<evidence>"
agentdb emit command "experiment-conclude" "" '{"H":"ID","EXP":"ID","verdict":"X","confidence":0.XX}'
```

Loop back to PICK for next hypothesis.
</phase> <phase id="judge" name="JUDGE — Review All Evidence" trigger="sense.sufficient_evidence"> For each hypothesis with >= 3 experiments: summarize evidence, calculate final confidence, classify graduated | refuted | needs-more-evidence | inconclusive.
```bash
agentdb hypothesis export
```
Write detailed report to _meta/research/experiment-report.md. Proceed to EVOLVE.
</phase> <phase id="evolve" name="EVOLVE — Self-Reconfigure"> The emergent part. The system reconfigures based on evidence.
**Graduate** (confidence >= 0.8, evidence_for >= 3, ratio >= 3:1):
promote via the artifact ladder with human approval — hook if enforceable, agent if
a role, skill if methodology; CLAUDE.md prose only as last resort.

**Kill** (confidence < 0.2, evidence_against >= 2):
propose rule removal from CLAUDE.md (present to human).

**Mutate** (inconclusive after 5+ experiments):
the hypothesis may be poorly stated; propose a refined version as a NEW hypothesis,
linked to the original (evolution chain).

<ask_user>
  Use AskUserQuestion ONCE at the end of the evolve phase:
  Ask: "{graduated} rules proven, {killed} rules disproven, {mutated} rules refined. Apply changes?"
  Options: apply all, review individually, skip for now
</ask_user>

```bash
agentdb emit command "experiment-evolve" "" '{"graduated":N,"killed":N,"mutated":N}'
```
</phase> </cycle>

<loop_control> continue_if: untested hypotheses remain OR new evidence changes confidence significantly pause_at: EVOLVE phase (only human checkpoint — graduation/kill decisions) stop_if: all hypotheses have >= 3 experiments AND no graduation/kill candidates on_stop: write final report to _meta/research/experiment-report.md, agentdb write-end Iteration budget: max 20 cycles per invocation. </loop_control>

<anti_patterns> <never>Confirm a hypothesis without running a real experiment.</never> <never>Use a single data point to graduate a hypothesis.</never> <never>Ignore refuting evidence because the rule "feels right".</never> <never>Test a hypothesis with a method that can only confirm (design for falsifiability).</never> <never>Modify the hypothesis after seeing results (that is a new hypothesis).</never> </anti_patterns>

<hard_stops>

  • NEVER modify CLAUDE.md autonomously. Present changes at EVOLVE, human decides.
  • NEVER delete hypotheses. Mark as refuted. Audit trail is sacred.
  • NEVER fabricate evidence. If experiment can't run, mark inconclusive with reason.
  • NEVER run destructive experiments without explicit approval.
  • ALWAYS record evidence, even for inconclusive results. </hard_stops>

<on_end>

agentdb write-end '{"skill":"experiment","cycles":N,"experiments_run":N,"graduated":N,"refuted":N,"mutated":N}'

</on_end>

</skill>

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

同じリポジトリのスキル

概要と使いどころ

Use AtomLane to compile and execute safe atomic parallel plans on macOS and native Windows Preview for worthwhile independent argv tasks, dependency DAGs, supported platform entrypoints, or Apple-silicon operators. Use at task start or an execution boundary when structured local work may contain two or more worthwhile units; skip plain answers, one quick command, and work whose effects cannot be safely bounded.

日本語の概要は準備中です。原文の説明を表示しています。

hashgraph-online/awesome-codex-plugins1,2732026年10月10日 更新

add

無料

Register a deferred decision in the debt registry. Trigger by judgment, not a marker scan, whenever a future reader would ask "why this way?": an unmade decision, stub, loosened type, bypassed check, swallowed error, a default picked "for now", or a TODO/FIXME/HACK/XXX marker. Trigger immediately whenever you defer work, or when the user invokes $add. Over-register freely; the developer drops with "drop A", "drop A,C", or "drop all".

日本語の概要は準備中です。原文の説明を表示しています。

hashgraph-online/awesome-codex-plugins1,2732026年10月10日 更新

ADK 框架适配层。为 LangChain / EINO / AutoGen / AgentScope / CrewAI 提供框架特定的 代码模板、惯用模式、API 映射和项目结构,供 agent-dev-workshop Phase 5 代码生成使用。 每个框架 reference 文件标注 verified_date 用于版本锁定。

日本語の概要は準備中です。原文の説明を表示しています。

hashgraph-online/awesome-codex-plugins1,2732026年10月10日 更新

中文调试修复技能。用于报错、测试失败、页面异常、功能不符合预期、需要定位根因并做最小修复时。触发语包括"进入调试模式""帮我修问题""报错了""测试失败""页面坏了""找根因"。

日本語の概要は準備中です。原文の説明を表示しています。

hashgraph-online/awesome-codex-plugins1,2732026年10月10日 更新

交互式 AI Agent 开发工作坊:通过 6 阶段深度协作对话,引导用户完成 Agent 需求分析、架构设计、 工具定义、Prompt 与编排设计、代码生成、验证迭代,产出可直接运行的 Agent 项目。 框架无关设计优先,支持 LangChain / EINO / AutoGen / AgentScope / CrewAI 等 ADK 框架。

日本語の概要は準備中です。原文の説明を表示しています。

hashgraph-online/awesome-codex-plugins1,2732026年10月10日 更新

中文漂移审计技能。用于项目或学习过程变乱、上下文漂移、任务分叉、多个方案冲突、命名不一致、Codex 可能顺手改多了时。触发语包括"漂移检查""感觉跑偏了""项目变乱了""检查是否失控""分叉太多""上下文漂移"。

日本語の概要は準備中です。原文の説明を表示しています。

hashgraph-online/awesome-codex-plugins1,2732026年10月10日 更新

hashgraph-online のスキルをすべて見る

このスキルの問題を報告する