# Experiment Plan: [feature / change]

> A/B test design – decide the success criteria before looking at the
> data, or you will rationalize whatever you see.

| Field | Value |
|-------|-------|
| Experiment | [name] |
| Owner | @name |
| Analyst | @name |
| Start | YYYY-MM-DD |
| Planned end | YYYY-MM-DD |
| Status | Design / Running / Decided |

## Hypothesis

We believe **[change]** will cause **[metric]** to **[direction]** by
**[size]** because **[mechanism]**.

## Variants

| Variant | Change | Allocation |
|---------|--------|------------|
| Control | current experience | N% |
| A | [change] | N% |

## Metrics

| Type | Metric | Direction | Threshold |
|------|--------|-----------|-----------|
| Primary | [e.g. activation rate] | up | >= +N% |
| Guardrail | [e.g. churn] | flat or better | <= +N% |
| Diagnostic | [e.g. click-through] | informational | – |

## Sizing

- Baseline conversion: N%
- Minimum detectable effect: N%
- Significance: 95%; power: 80%
- Required sample per arm: $n = \frac{16\sigma^2}{\delta^2}$ = N
- Daily traffic: N -> **expected duration: N days**

Commit to the duration now. Peeking early and stopping on a spike is
the most common way experiments lie.

## Segmentation & exclusions

- Included: [population]
- Excluded: [internal users, bots, ...]
- Pre-declared subgroups to analyze: [list]

## Decision rule

| Outcome | Action |
|---------|--------|
| Primary up, guardrails flat | Ship variant |
| Primary flat | Revert; record learning |
| Guardrail broken | Stop immediately, revert |

## Results (fill after)

| Variant | Primary metric | Guardrail | Verdict |
|---------|----------------|-----------|---------|
| Control | | | |
| A | | | |

## Learnings

What this changes in the roadmap:
