本文へ移動
cccskills
無料GitHub で公開

experimentation

Designs, runs, and reads A/B tests and growth experiments — hypothesis, sample size, duration, and honest interpretation. Use this to plan a test, judge whether a result is real, build an experimentation program, decide what to test next, or diagnose why tests keep producing inconclusive or non-replicating results.

インストール方法を見る

含まれるファイル(2)

  • SKILL.md3.8 KB
  • references/sources.md1.4 KB

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

Experimentation

Most A/B testing programs produce confident conclusions from insufficient data. The discipline is almost entirely in what you do before launch.

Before running

  • Hypothesis with a mechanism. "Moving the pricing table above the fold will raise trial starts, because visitors currently leave before seeing pricing." Not "let's try a green button."
  • One primary metric, chosen in advance. Secondary metrics are context, never the verdict.
  • Sample size calculated in advance, from your baseline rate and the smallest lift that would change a decision. If the required sample is unreachable, do not run the test — decide by judgment and say so.
  • Duration set in advance, covering at least one full weekly cycle, and two if the buying cycle is long.
  • Guardrail metrics that would make you reject a win: refunds, support volume, downstream retention.

While running

Do not look at results and act on them mid-flight. Peeking and stopping at significance is the single most common way to generate false positives, and it is very effective at it.

Check only that the test is running correctly — even split, no broken variant, tracking firing.

Reading

  • At the pre-set duration, not before, and not extended because it is nearly significant. Extending until significance manufactures it.
  • Significance is not size. A statistically significant 0.3% lift may not be worth shipping.
  • Inconclusive is a real result and the most common one. It means the change did not matter enough to detect, which is useful.
  • Check the guardrails before declaring a win.
  • Segment afterward for hypotheses only, never for verdicts. Slice enough ways and something is always significant.

Program level

Test where the traffic and the leverage are. Most sites can only run a handful of adequately powered tests a year — spend them on structural questions, not button colors.

Keep a log of every test: hypothesis, result, decision. Without it, teams re-run the same tests every eighteen months and re-learn the same things.

Sources

references/sources.md in this skill lists the outside authorities that settle the questions here — what each one is authoritative for, and what you may do with it. Check them before answering on anything they cover, and cite what you used. Most are free to read and not free to reproduce; the use note on each is binding.

Tooling

Client-side and web testing: Optimizely, VWO, AB Tasty, and similar. Warehouse- or product-native: GrowthBook, Statsig, Eppo, PostHog, and similar — these compute against your own event data, which is what you want once the metric definitions matter.

Feature flags are the server-side path to the same thing: LaunchDarkly, Unleash, Split, and similar. Running an experiment behind a flag you already use for release control is cheaper than adding a second system, and technology:release-and-deployment covers the release side of it.

No tool fixes an underpowered test. The platform reports a result either way, which is exactly the risk.

Never

  • Stop a test because it reached significance early. Peeking until it looks conclusive manufactures the result.
  • Run a test that cannot reach adequate sample size in a reasonable window. Ship the change on judgment instead and say so.
  • Change more than one variable and attribute the outcome to the one you liked.
  • Count a flat result as a failure. A well-run test that rules out a plausible idea has bought information.

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

同じリポジトリのスキル

概要と使いどころ

Designs and audits who can reach what — authentication, authorization models, privileged access, service credentials, and joiner-mover-leaver process. Use this to design a permissions model, run an access review, reduce standing privilege, handle offboarding, set up SSO or MFA, manage service and machine credentials, or diagnose why permissions have sprawled.

日本語の概要は準備中です。原文の説明を表示しています。

cbrock84/headcount2,0312026年9月18日 更新

Concentrates marketing and sales effort on a named set of accounts rather than on volume — qualifying whether the model fits your economics at all, building the account list and the buying group inside each, tiering effort against account value, coordinating so the account experiences one campaign rather than several, and measuring account progression instead of leads. Use this to decide whether to run an account-based program, build one, or work out why an existing one produces activity and no pipeline.

日本語の概要は準備中です。原文の説明を表示しています。

cbrock84/headcount2,0312026年9月18日 更新

Gets new users from signup to first real value — signup flow, onboarding, time-to-value, and the early experience that determines whether someone becomes a user or a lapsed account. Use this to design or fix signup and onboarding, diagnose why signups do not convert to active use, reduce time-to-value, or decide what a new user must accomplish first.

日本語の概要は準備中です。原文の説明を表示しています。

cbrock84/headcount2,0312026年9月18日 更新

Designs orchestrator-and-subagent hierarchies for a repository — splitting agents by exclusive write surface, pairing every producer with an independent auditor, and enforcing the split with a script that runs in CI. Use this whenever the user wants to set up, expand, audit, or fix a multi-agent or subagent structure for a codebase; asks how to divide work between agents; wants agent charters, roles, or a surface map written; or is hitting agents that collide on the same files, review their own work, or drift from their remit. Also use when sizing a roster or deciding whether a new agent is justified.

日本語の概要は準備中です。原文の説明を表示しています。

cbrock84/headcount2,0312026年9月18日 更新

Governs models and AI systems in production — intended use, evaluation, monitoring, human oversight, documentation, and the decision to deploy or retire. Use this before deploying a model or AI feature, when defining evaluation criteria, when a model's behavior has drifted, when assessing AI risk or regulatory exposure, or when deciding whether an AI system is fit for a consequential decision.

日本語の概要は準備中です。原文の説明を表示しています。

cbrock84/headcount2,0312026年9月18日 更新

Produces executive-level research — market sizing, competitor mapping, trend analysis, and strategic intelligence — grounded in cited sources with the confidence in each claim made explicit. Use this to analyze a market or industry, map competitors, evaluate a market-entry or build-versus-buy decision, produce a research brief, or assemble evidence for a decision. Also use when comparing options that need a structured, evidence-based verdict rather than an opinion.

日本語の概要は準備中です。原文の説明を表示しています。

cbrock84/headcount2,0312026年9月18日 更新

cbrock84 のスキルをすべて見る

このスキルの問題を報告する