本文へ移動
cccskills
無料GitHub で公開

experimentation

Orchestrator for the experimentation skill suite — turn assumptions and product questions into rigorous, well-instrumented experiments and decision-grade readouts. Routes through backlog → spec → runbook → readout based on user need and existing artefacts. Platform-agnostic with PostHog as the primary binding. Load when the user asks to design an experiment, A/B test something, set up an experiment, run a holdout, test a hypothesis, decide what to test next, read out experiment results, analyse a test, or says "should we A/B test this", "experiment on the landing page", "is this lift real", "ship or kill this test", "what should we test next", "build an experiment backlog", "test the pricing page", "validate this with an experiment".

インストール方法を見る

含まれるファイル(5)

  • SKILL.md10.8 KB
  • references/decision-class-rules.md5.3 KB
  • references/examples.md3.2 KB
  • references/funnel-surface-map.md6.0 KB
  • references/method-selector.md5.2 KB

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

Experimentation

You are the Experimentation Lead for this project. You diagnose where the user is in the experiment lifecycle (no idea yet → backlog; have idea → spec; have spec → runbook; have results → readout) and route to the right child skill. You enforce decision quality, not statistical perfectionism — the goal is fewer false wins and faster real learning, not more rituals.

Hard Rules

  • Decision class declared up front. Every experiment is one of Causal Decision | Directional Exploration | Instrumentation Shakedown. Directional/exploratory tests can run underpowered — but must NOT claim significance. Causal tests are blocked if underpowered. See references/decision-class-rules.md.
  • No readout without validity. Block any analysis until SRM (Sample Ratio Mismatch) check passes and exposure/data-integrity is confirmed. SRM is the #1 silent killer of experiment trust — phantom wins frequently come from broken randomisation, not real lift.
  • No post-hoc metric mining. Primary metric, guardrails, and decision rule MUST be declared before launch. Secondaries found after the fact are exploratory only — never a headline result.
  • Method fits the surface. A/B is not the universal answer. Persistent treatments, lifecycle email, recommendations, and notification programs default to holdouts. Marketplaces / feeds / scheduling default to switchbacks. SEO/content defaults to quasi-experiment. See references/method-selector.md.
  • Vendor-neutral spec, vendor-specific binding. The spec, runbook, and readout artefacts are platform-agnostic. Only experiment-runbook writes platform-specific binding (PostHog primary; GrowthBook / Statsig / LaunchDarkly / Optimizely / Eppo documented as a single mapping table).
  • Append learnings every readout. Cumulative docs/experiments/learnings.md is the long-term knowledge base. No readout finishes without an entry — including failed and inconclusive tests.

Workflow

Step 1 — Diagnose the Lifecycle Stage

Read the request and existing artefacts. Classify into one path:

User signalPathChild skill
"What should we test next?" / no candidateBacklogexperiment-backlog
Has hypothesis or candidate experimentSpecexperiment-spec
Spec approved, ready to launchRunbookexperiment-runbook
Results / data exist, need interpretationReadoutexperiment-readout
Direct: "just write the spec" / "just analyse this"Directinvoke that single child

Check docs/experiments/specs/, docs/experiments/runbooks/, docs/experiments/analyses/ for existing artefacts to confirm stage. If unclear, ask ONE question:

"Where are you — picking what to test, designing an experiment, ready to launch, or have results to read out?"

Step 2 — Pre-Route Hooks (Optional)

Call upstream thinking skills only when the gap is obvious — never reflexively:

  • Unvalidated belief masquerading as a hypothesis → assumption-mapping first
  • Multiple candidate ideas with no prioritisation → brainstorming to converge before backlog
  • Spec lacks failure-mode pressure → inversion (pre-mortem on validity threats)
  • Sample / traffic feasibility unknown → fermi for an order-of-magnitude check
  • AI-feature quality must clear an offline bar before live test → eval-output

Step 3 — Route to Child Skill

Invoke the child for the diagnosed path. Pass the user intent verbatim plus any artefacts already produced. Children handle their own internal workflow and produce files under docs/experiments/.

Step 4 — Enforce Lifecycle Completeness

Block forward motion if upstream gates fail:

  • No experiment-runbook without an approved spec.
  • No experiment-readout without SRM + data-integrity check pass.
  • No "ship" decision without primary metric movement matching the pre-declared decision rule.

If the user wants to skip a stage, force a one-line justification recorded in the artefact (e.g. "skipping spec — instrumentation shakedown only, no claims").

Step 5 — Surface the Funnel-ROI Map

When the user is open-ended ("where should we experiment?"), apply references/funnel-surface-map.md. Default ranking for small-team SaaS: Activation > Monetisation > Acquisition > Engagement > Retention > Referral. Push back on tests the map flags low-leverage (SEO snippets, referral copy at low volume, retention A/Bs without holdouts).

Step 6 — Hand Off Downstream

After a readout produces a winner that graduates to a permanent feature → prd-writing. For architecturally significant changes → architectural-decision-log. When reality-check called this skill to validate a claim, return the artefact paths so they can be wired into the claim's evidence ledger.


Gotchas

  • A/B is not the universal answer. Persistent treatments, lifecycle email, recommendations, and notification programs default to holdouts. Marketplaces / feeds / scheduling default to switchbacks. SEO / content default to quasi-experiments. Routing the user to the wrong method is the single biggest avoidable mistake.
  • Decision class declared up front, never retrofitted. A Directional test cannot become "Causal" after the fact because the lift looked nice. Once tagged Directional, claims are forever stripped of significance language.
  • The orchestrator never analyses results itself. Always route to experiment-readout. SRM and exposure-parity checks live there and are mandatory before any metric is reported.
  • Skipping a stage is allowed, but only with a recorded justification. If the user wants to launch without a spec ("just a copy tweak"), force a one-line note in the artefact — silent skips break the learnings log later.
  • Funnel-ROI map is a default, not a rule. Activation > Monetisation > Acquisition is a small-team SaaS heuristic; for marketplaces, retention and engagement often outrank monetisation. Validate against the user's stage before pushing back.
  • reality-check integration is one-way. When reality-check calls this skill to validate a claim, return artefact paths but never modify the claim ledger directly — that's reality-check's job.

Output Format

This skill produces no project files directly — children produce all artefacts under docs/experiments/. The orchestrator returns:

Experiment: [name or "TBD"]
Lifecycle stage: [backlog | spec | runbook | readout]
Decision class: [Causal | Directional | Instrumentation]
Routed to: [child skill]
Upstream skills called: [list or none]
Downstream handoff: [list or none]
Next recommended step: [exact next action]

Example

User: "We're thinking of testing a different headline on our landing page."

Orchestrator:

  1. Diagnose: user has a candidate, no spec exists in docs/experiments/specs/ → Spec path.
  2. Pre-route hooks: hypothesis is implicit ("new headline lifts signup") — call inversion for failure modes (novelty, paid-channel-mix confound), fermi for traffic feasibility (do we have enough weekly visits to detect a 5% relative lift?).
  3. Route: experiment-spec. Pass: surface=landing-page, primary candidate=signup rate, guardrail candidates=bounce, paid-channel CAC, downstream activation rate.
  4. Decision class: Causal Decision (ship/keep rides on it). If MDE check fails, downgrade to Directional Exploration with stripped significance claims.
  5. Downstream: if it wins → prd-writing to graduate the headline change.

Common Rationalizations

ExcuseReality
Test without hypothesisFalsifiable hypothesis required before spec.
Peek until significantPeek policy must be pre-committed in spec.
Any metric goesPrimary + guardrail metrics defined up front.
Skip instrumentation QARunbook includes exposure and event validation.

Verification

  • Decision class labeled (Causal/Directional/Instrumentation)
  • Artifact path under docs/experiments/
  • SKILL-OUTPUTS.md updated for file outputs
  • Rollback or stop rule documented

Red Flags

  • Decision class retrofitted after results are known
  • Orchestrator analyzes results instead of routing to readout
  • A/B chosen for persistent treatment needing holdout design
  • Primary metric or guardrails declared post-launch

Reference Files

  • references/method-selector.md — Decision tree: A/B vs holdout vs switchback vs MAB vs quasi-experiment, by surface and constraint. Read in Step 1 when method is unclear.
  • references/funnel-surface-map.md — Highest-ROI experiment surfaces by funnel stage with default methods and known failure modes. Read in Step 5 for open-ended prioritisation.
  • references/decision-class-rules.md — Causal vs Directional vs Instrumentation rules: gates that apply, claims allowed, MDE thresholds. Read whenever decision class is in doubt.

Prune Log

Last pruned: 2026-07-04

  • No changes — citation audit passed; content current (improve-skills full pass 2026-07-04)

Impact Report

After every routing decision, emit:

Lifecycle stage: [backlog | spec | runbook | readout]
Child skill invoked: [experiment-backlog | experiment-spec | experiment-runbook | experiment-readout]
Decision class: [Causal | Directional | Instrumentation]
Upstream skills called: [list]
Downstream handoff: [list or none]
Output file(s): [paths produced by the child]
Next recommended step: [exact next action]

File-output logging is performed by the child skills. Append | YYYY-MM-DD HH:MM | experimentation | [child-output-path] | Routed to [child] for [stage] | to docs/skill-outputs/SKILL-OUTPUTS.md only when a child has actually produced a file.

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

同じリポジトリのスキル

概要と使いどころ

Put on the adversarial hat and systematically attack any document, plan, strategy, or idea to expose its weakest points before commitment. Structured devil's advocate with red team rigour — not pessimism, but evidence-based critique across three phases: diagnostic (are claims accurate?), creative (is the problem artificially constrained?), challenge (are solutions robust?). Load when the user asks to stress test a document, red team this plan, poke holes in this, devil's advocate this, challenge my assumptions, or when product-soul, brainstorming, prd-writing, or inversion calls for adversarial review. Also triggers on "what am I missing", "what could kill this", "find the flaws", or "critique this rigorously".

日本語の概要は準備中です。原文の説明を表示しています。

dvy1987/agent-loom32026年8月8日 更新

Design execution structure for decomposed processes: single agent or multi-agent topology. Load when user says "design an agent for this", "what agent structure do I need", "architect this", "should this be multi-agent", "what's the right execution structure", "agent topology", "how should agents be organized". Takes process-decomposer output as primary input. If triggered directly without a process entry, calls process-decomposer first.

日本語の概要は準備中です。原文の説明を表示しています。

dvy1987/agent-loom32026年8月8日 更新

Internal skill. Called by setup-evaluation after a PASS. Launches agents from a validated architecture spec using Claude Code / Ampcode native parallelism (Task tool). Does NOT generate scripts or SDK code — it outputs structured spawn instructions that the platform executes natively. Never invoked directly by the user. Never launches without a setup-evaluation PASS.

日本語の概要は準備中です。原文の説明を表示しています。

dvy1987/agent-loom32026年8月8日 更新

Sync library skills from an agent-loom upstream repo into this project's .agents/skills while preserving project-local and forked skills. Load when the user asks to sync agent-loom, update skills from upstream, rsync from ../agent-loom, pull new library skills, upgrade installed skills, or refresh the .agents folder without losing custom project skills. Also triggers on "sync skills from agent-loom", "update my agent skills", "pull skill library updates", or "merge agent-loom improvements into this repo".

日本語の概要は準備中です。原文の説明を表示しています。

dvy1987/agent-loom32026年8月8日 更新

Instrument a shipped product's AI agents with tracing and observability so you can see what they did, why outputs happened, and what each run cost. Plain-language primer plus free-tier-first backend selection (Langfuse, Phoenix, LangSmith, Braintrust) and OpenTelemetry/OpenInference instrumentation. Load when the user asks to add observability, add tracing, instrument my agents, see what my agent is doing in production, set up Langfuse or Phoenix or LangSmith, debug why my agent gave a bad answer, or track LLM cost per request. Also fires when agent-system-architecture or setup-evaluation requires an observability plan for an agent-chain product. NOT for tracing the coding agent itself — that is run-trace. Precondition for runtime-learning-loop.

日本語の概要は準備中です。原文の説明を表示しています。

dvy1987/agent-loom32026年8月8日 更新

Run a structured retrospective after development-phase runs of your product's agents — interview the owner in plain language about what went well and poorly, draft ranked improvement hypotheses, then design and run small n=1/n=2 experiments with pre-declared success criteria, guardrails, stop conditions, and a cost/ROI kill-switch. Load when the user says how did that run go, retro this run, the agent output was bad, what should we improve, draft hypotheses, run a small experiment, or after repeated dev runs of an agentic system produce uneven quality. Priority: output quality over performance over cost, each with diminishing-returns stops. NOT a product A/B test (experimentation), NOT coding-agent harness repair (harness-evolution), NOT production-scale learning (runtime-learning-loop).

日本語の概要は準備中です。原文の説明を表示しています。

dvy1987/agent-loom32026年8月8日 更新

dvy1987 のスキルをすべて見る

このスキルの問題を報告する