本文へ移動
cccskills
無料GitHub で公開

codexkit-a-b-test-planner

Design rigorous A/B test plans with hypothesis, sample size calculation, Minimum Detectable Effect (MDE), randomization strategy, and decision rules. Includes guardrail metrics and rollout playbook. Use when planning product experiments, conversion optimization, or data-driven feature decisions.

インストール方法を見る

含まれるファイル(5)

  • SKILL.md5.5 KB
  • agents/openai.yaml165 B
  • examples/common-mistakes.md794 B
  • examples/good-output.md870 B
  • verification/checklist.md1.4 KB

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

A/B Test Planner

When to Use

  • Before running any product experiment or feature test
  • When optimizing conversion funnels or UX flows
  • When leadership requires statistical rigor for feature decisions
  • When planning multi-variant tests or sequential experiments

Procedure

Step 1 — Hypothesis

Write a clear, falsifiable hypothesis:

  • If [change we're making]
  • Then [metric we expect to change]
  • Because [reasoning / user insight]

Step 2 — Metrics

TypeMetricCurrent Baseline
PrimaryThe metric that determines success[value]
SecondarySupporting metrics that provide context[value]
GuardrailMetrics that must NOT degrade[value]

Step 3 — Sample Size & Duration

Calculate required sample size using:

  • Baseline conversion rate (p₁)
  • Minimum Detectable Effect (MDE) — smallest meaningful change
  • Statistical significance level (α) — typically 0.05
  • Statistical power (1−β) — typically 0.80

n = f(p₁, MDE, α, β) → use standard sample size calculator

Duration = n / (daily traffic × allocation %)

Step 4 — Randomization Plan

  1. Randomization unit: user, session, device, or account
  2. Allocation: 50/50, or asymmetric with justification
  3. Stratification: any segments to balance (geography, plan, device)
  4. Exclusion: users to exclude (employees, bots, existing tests)

Step 5 — Decision Rules

OutcomeCriteriaAction
WinnerPrimary metric ↑ ≥ MDE, p < 0.05, guardrails stableShip to 100%
NeutralNo significant differenceKeep control, iterate hypothesis
LoserPrimary metric ↓ significantlyRevert, analyze why
Guardrail breachAny guardrail metric degrades > thresholdStop test immediately

Step 6 — Rollout Playbook

  1. Ramp: 5% → 25% → 50% → 100% over [days]
  2. Monitoring: check metrics daily during ramp
  3. Rollback trigger: guardrail breach or unexpected anomaly

Inputs

InputRequiredFormat
Change descriptionYesWhat is being tested
Baseline metricYesCurrent value
MDE targetYesPercentage or absolute
Daily trafficYesNumber of users/events
Test duration budgetRecommendedMax days willing to run

Output

## A/B Test Plan — [Test Name]

### Hypothesis
If we simplify the checkout form from 5 fields to 3 fields,
then checkout completion rate will increase by ≥ 5%,
because user research shows 40% abandon at the address step.

### Metrics

| Type | Metric | Baseline | Target |
|------|--------|----------|--------|
| Primary | Checkout completion rate | 45% | ≥ 50% |
| Secondary | Average order value | $65 | Stable |
| Guardrail | Revenue per user | $12 | No decrease |
| Guardrail | Error rate | 0.5% | No increase |

### Sample Size

| Parameter | Value |
|-----------|-------|
| Baseline rate | 45% |
| MDE | 5% (absolute) |
| Significance (α) | 0.05 |
| Power (1−β) | 0.80 |
| Required n per variant | ~1,600 |
| Daily traffic | 800 users |
| Allocation | 50/50 |
| **Estimated duration** | **4 days** |

### Randomization
- Unit: User (cookie-based)
- Allocation: 50% control / 50% variant
- Exclusion: Internal users, users in other active tests

### Decision Rules
[As defined in procedure]

### Rollout Playbook
Day 1–2: 10% ramp → monitor → Day 3–4: 50% → Day 5: 100%

Definition of Done

  • Hypothesis is clear and falsifiable
  • Primary, secondary, and guardrail metrics defined
  • Sample size calculated with stated parameters
  • Duration estimated based on traffic
  • Randomization plan documented
  • Decision rules and rollout playbook included

Examples

Prompt

We want to test if a new onboarding flow increases activation rate.
Current activation: 32%. MDE: 3 percentage points. Daily new users: 500.
We have max 14 days for the test.
Design a complete A/B test plan with sample size and decision rules.

Quality Criteria

  • Every finding is tied to a specific evidence source (log, test, metric)
  • Pass/fail criteria are binary and measurable — no subjective judgments
  • Severity levels are assigned with clear thresholds
  • Remediation steps are provided for all critical and high findings

Verification (4C)

CheckQuestion
CorrectnessAre all pass/fail criteria applied against the correct standard or rule?
CompletenessWere all required dimensions or checklist items evaluated?
Context-fitDoes the verification scope match the actual risk level of the deliverable?
ConsequenceIf this passed verification but had a hidden flaw, what is the worst-case impact?

Edge Cases

  • Incomplete data for full assessment — Document which checks were limited and flag for re-verification when data becomes available.
  • Ambiguous pass/fail criteria — Request clarification from the standard owner before scoring. Mark as 'Needs Review'.
  • Multiple overlapping standards — Identify the governing standard and note where others diverge.

Changelog

  • v1.0.0 — Initial release

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

同じリポジトリのスキル

概要と使いどころ

Review REST and GraphQL API designs for consistency, usability, and best practices. Covers naming conventions, versioning strategy, error format, pagination, authentication patterns, and breaking change detection. Use when reviewing API specs, designing new APIs, or auditing existing endpoints.

日本語の概要は準備中です。原文の説明を表示しています。

hoavdc/CodexKit252026年10月8日 更新

Write Architecture Decision Records (ADRs) following the Michael Nygard format. Captures context, options considered, decision rationale, and consequences. Use when making technology choices, framework selections, or any architectural decision that future developers need to understand.

日本語の概要は準備中です。原文の説明を表示しています。

hoavdc/CodexKit252026年10月8日 更新

Assess organizational readiness for financial audits (internal or external). Map assertions to account balances, check evidence completeness, score readiness using a Red/Amber/Green framework, and generate a remediation timeline. Aligned with SOX, IFRS, and GAAP audit standards. Use before scheduled audits or when preparing for first-time compliance.

日本語の概要は準備中です。原文の説明を表示しています。

hoavdc/CodexKit252026年10月8日 更新

Design safe recurring Codex automations with clear prompts, outputs, schedules, and gating rules.

日本語の概要は準備中です。原文の説明を表示しています。

hoavdc/CodexKit252026年10月8日 更新

Refine Product Backlog Items to meet INVEST criteria. Write User Stories with Acceptance Criteria in Given/When/Then format, estimate with Story Points, and flag dependencies. Use before sprint planning when backlog items need grooming. Do not use to prioritize the backlog — that is the Product Owner's decision.

日本語の概要は準備中です。原文の説明を表示しています。

hoavdc/CodexKit252026年10月8日 更新

Build or refine brand positioning with audience, category, differentiators, proof, tone, JTBD signals, and competitive context. Use when marketing, founders, or GTM teams need a positioning canvas, messaging pillars, or campaign foundation. Do not use for isolated ad copy tweaks with no strategy question.

日本語の概要は準備中です。原文の説明を表示しています。

hoavdc/CodexKit252026年10月8日 更新

hoavdc のスキルをすべて見る

このスキルの問題を報告する