A/B Testing Cheat Sheet
A practical guide to designing, running, and statistically analyzing A/B tests, covering sample size, significance testing, and common pitfalls.
Sample Size & Z-Test
Power analysis and two-proportion z-test.
import numpy as npfrom statsmodels.stats.power import NormalIndPowerfrom statsmodels.stats.proportion import proportions_ztest# Sample size needed to detect an effect (power analysis)analysis = NormalIndPower()n_per_group = analysis.solve_power( effect_size=0.2, # standardized effect size (Cohen's h) alpha=0.05, # significance level power=0.8, # desired statistical power ratio=1.0,)print(f"Required sample size per group: {n_per_group:.0f}")# Two-proportion z-test for conversion rate comparisonconversions = np.array([120, 150]) # [control, treatment]visitors = np.array([2000, 2000])z_stat, p_value = proportions_ztest(conversions, visitors)print(f"z = {z_stat:.3f}, p = {p_value:.4f}")
T-Test for Continuous Metrics
Compare means and build a confidence interval.
from scipy import statsimport numpy as np# For continuous metrics (e.g. revenue per user), use a t-testcontrol = np.random.normal(50, 10, 1000)treatment = np.random.normal(52, 10, 1000)t_stat, p_value = stats.ttest_ind(control, treatment, equal_var=False)print(f"t = {t_stat:.3f}, p = {p_value:.4f}")# 95% confidence interval for the difference in meansdiff = treatment.mean() - control.mean()se = np.sqrt(control.var(ddof=1) / len(control) + treatment.var(ddof=1) / len(treatment))ci = (diff - 1.96 * se, diff + 1.96 * se)print(f"Difference: {diff:.2f}, 95% CI: {ci}")
A/B Testing Concepts
Key vocabulary for designing experiments.
- Null hypothesis (H0)- assumes no difference between control and treatment
- Statistical significance (p-value)- probability of seeing this result (or more extreme) if H0 were true
- Statistical power- probability of correctly detecting a true effect (typically target 80%)
- Minimum Detectable Effect (MDE)- smallest effect size the test is designed to reliably detect
- Sample Ratio Mismatch (SRM)- unequal group sizes vs. the intended split, signals a bug in randomization
- Novelty effect- short-term behavior change from users noticing something new, not a lasting effect
- Peeking problem- repeatedly checking significance before the planned sample size inflates false positive rate
- Multiple comparisons- testing many metrics/segments increases chance of a false positive; use a correction (e.g. Bonferroni)
Test Design Checklist
Steps to run a trustworthy experiment.
- Define one primary metric- pick a single primary success metric before launch to avoid p-hacking
- Compute required sample size- run a power analysis before starting, don't stop early
- Randomize consistently- assign users to variants deterministically (e.g. hash of user ID)
- Check for SRM- verify observed group sizes match the intended split before trusting results
- Run for full business cycles- avoid day-of-week effects by running at least 1-2 full weeks
- Pre-register the analysis plan- decide the test and segments in advance to avoid fishing for significance
Bayesian A/B Test with Beta-Binomial
Compute the probability that treatment beats control using conjugate priors, no p-values needed.
import numpy as np# Beta(1,1) = uniform prior; update with observed conversions/non-conversionscontrol_conversions, control_visitors = 120, 2000treat_conversions, treat_visitors = 150, 2000alpha_c, beta_c = 1 + control_conversions, 1 + (control_visitors - control_conversions)alpha_t, beta_t = 1 + treat_conversions, 1 + (treat_visitors - treat_conversions)# Monte Carlo: sample from each posterior and comparerng = np.random.RandomState(0)samples_c = rng.beta(alpha_c, beta_c, 100_000)samples_t = rng.beta(alpha_t, beta_t, 100_000)prob_treat_better = (samples_t > samples_c).mean()expected_uplift = (samples_t / samples_c - 1).mean()print(f"P(treatment > control): {prob_treat_better:.3f}")print(f"Expected relative uplift: {expected_uplift:.2%}")
CUPED Variance Reduction
Use pre-experiment data to shrink metric variance and detect smaller effects with the same sample size.
import numpy as np# y = experiment metric, x = same metric measured pre-experiment (covariate)def cuped_adjust(y: np.ndarray, x: np.ndarray) -> np.ndarray: theta = np.cov(x, y)[0, 1] / np.var(x) # regression coefficient return y - theta * (x - x.mean())y_control = np.random.normal(50, 10, 2000)x_control = y_control + np.random.normal(0, 3, 2000) # pre-period correlated metricy_adjusted = cuped_adjust(y_control, x_control)print(f"Raw variance: {y_control.var():.2f}")print(f"CUPED-adjusted variance: {y_adjusted.var():.2f}")# lower variance => narrower confidence intervals => higher effective power
Sequential Testing (Group Sequential Bound)
Check significance mid-experiment without inflating false positive rate, using an O'Brien-Fleming style spending function.
from scipy import statsimport numpy as npdef obrien_fleming_bound(alpha: float, info_fraction: float) -> float: """Return the z-critical-value threshold at a given fraction of planned sample size.""" z_alpha = stats.norm.ppf(1 - alpha / 2) # bound is very strict early on, relaxes toward the fixed-sample bound at fraction=1 return z_alpha / np.sqrt(info_fraction)# example: 3 planned interim looks at 33%, 66%, 100% of target sample sizefor frac in [0.33, 0.66, 1.0]: bound = obrien_fleming_bound(alpha=0.05, info_fraction=frac) print(f"At {frac:.0%} of sample: reject only if |z| > {bound:.3f}")
Bootstrap CI for Ratio Metrics
Get a valid confidence interval for ratio metrics (e.g. revenue per session) where the delta method is awkward.
import numpy as npdef bootstrap_ratio_ci(numerator: np.ndarray, denominator: np.ndarray, n_boot=5000, alpha=0.05): n = len(numerator) rng = np.random.RandomState(1) ratios = np.empty(n_boot) for i in range(n_boot): idx = rng.randint(0, n, n) ratios[i] = numerator[idx].sum() / denominator[idx].sum() lo, hi = np.percentile(ratios, [100 * alpha / 2, 100 * (1 - alpha / 2)]) return numerator.sum() / denominator.sum(), (lo, hi)# e.g. revenue per session, where sessions per user varies (violates i.i.d. assumption)revenue = np.random.exponential(20, 3000)sessions = np.random.poisson(2, 3000) + 1point_estimate, ci = bootstrap_ratio_ci(revenue, sessions)print(f"Revenue/session: {point_estimate:.3f}, 95% CI: {ci}")
Advanced Experimentation Concepts
Techniques and pitfalls beyond a basic fixed-horizon significance test.
- CUPED- Controlled-experiment Using Pre-Experiment Data; regresses out variance explained by a pre-period covariate to shrink confidence intervals
- Sequential testing / always-valid p-values- statistical methods (mSPRT, group sequential bounds) that let you peek at results continuously without inflating Type I error
- Network effects / interference- when treatment for one user affects control users (e.g. marketplaces, social feeds); violates SUTVA and biases results, mitigate with cluster or switchback randomization
- Simpson's paradox- an effect that reverses direction when segments are combined vs. viewed separately; always check for it before trusting an aggregate result
- Guardrail metrics- secondary metrics (latency, unsubscribe rate) monitored to catch harm even when the primary metric improves
- Heterogeneous treatment effects- the treatment effect varies by user segment; average effect can mask large gains for one group and losses for another
- Benjamini-Hochberg correction- controls the false discovery rate when testing many metrics/segments simultaneously, less conservative than Bonferroni
Decide your sample size and test duration up front using a power analysis, and don't stop the test early just because you see a significant p-value — 'peeking' at results repeatedly inflates your false positive rate far above 5%.