Probability Distributions Cheat Sheet
Reference for the most common discrete and continuous probability distributions, their parameters, and how to work with them in Python using scipy.stats.
Working with Distributions
PDF, CDF, PMF, and sampling with scipy.stats.
from scipy import stats# Normal (Gaussian) distributionnorm = stats.norm(loc=0, scale=1) # mean=0, std=1print(norm.pdf(0)) # probability density at x=0print(norm.cdf(1.96)) # P(X <= 1.96)samples = norm.rvs(size=1000) # random samples# Binomial distributionbinom = stats.binom(n=10, p=0.5) # 10 trials, p=0.5 successprint(binom.pmf(5)) # P(exactly 5 successes)# Poisson distributionpoisson = stats.poisson(mu=3) # average rate = 3 eventsprint(poisson.pmf(2)) # P(exactly 2 events)# Uniform distributionuniform = stats.uniform(loc=0, scale=10) # range [0, 10]
Fitting & Goodness-of-Fit
Fit a distribution to data and test normality.
from scipy import statsdata = [23, 25, 21, 22, 24, 20, 26, 27, 19, 24]# Fit a normal distribution to data (maximum likelihood estimate)mu, sigma = stats.norm.fit(data)# Test if data plausibly comes from a normal distributionstat, p_value = stats.shapiro(data) # Shapiro-Wilk normality testprint(f"p = {p_value:.4f}") # p < 0.05 suggests non-normal# Kolmogorov-Smirnov test against a specific fitted distributionks_stat, p_value = stats.kstest(data, "norm", args=(mu, sigma))
Distribution Reference
When to use each common distribution.
- Normal (Gaussian)- continuous, symmetric bell curve; parameters mean μ and std σ; models measurement errors, heights
- Binomial- discrete count of successes in n independent trials with success probability p
- Bernoulli- special case of binomial with n=1; a single yes/no trial
- Poisson- discrete count of events in a fixed interval given average rate λ; models rare, independent events
- Exponential- continuous time between events in a Poisson process; has the memoryless property
- Uniform- all outcomes in a range are equally likely
- Chi-square- distribution of the sum of squared standard normals; used in hypothesis tests
- Student's t- like normal but heavier tails; used for small-sample inference with unknown population std
Key Properties
Statistics used to summarize any distribution.
- Mean (expected value)- the long-run average outcome of the distribution
- Variance / standard deviation- measures spread around the mean
- Skewness- measures asymmetry; positive skew has a long right tail
- Kurtosis- measures tail heaviness relative to a normal distribution
- PDF vs PMF- PDF describes continuous distributions (density), PMF describes discrete ones (exact probability)
- CDF- cumulative distribution function; gives P(X <= x) for any x
Multivariate & Joint Distributions
Model correlated variables with a covariance matrix.
import numpy as npfrom scipy import statsmean = [0, 0]cov = [[1.0, 0.6], [0.6, 1.0]] # positive correlation between X and Ymvn = stats.multivariate_normal(mean=mean, cov=cov)sample = mvn.rvs(size=1000, random_state=0)print(mvn.pdf([0.5, 0.5])) # joint density at a point# Conditional distribution of Y given X=1 (still Gaussian for the MVN family)x_val = 1.0cond_mean = mean[1] + cov[1][0] / cov[0][0] * (x_val - mean[0])cond_var = cov[1][1] - cov[1][0] ** 2 / cov[0][0]cond_dist = stats.norm(loc=cond_mean, scale=np.sqrt(cond_var))# Copula-style dependence: transform correlated normals to arbitrary marginalsu = stats.norm.cdf(sample[:, 0]) # uniform marginal via probability integral transformexp_marginal = stats.expon.ppf(u, scale=2) # now Exponential, correlation structure preserved
Comparing Candidate Distributions
Fit several distributions and pick the best via AIC and KS tests.
import numpy as npfrom scipy import statsdata = np.random.default_rng(0).gamma(shape=2.0, scale=1.5, size=500)candidates = ["norm", "lognorm", "gamma", "weibull_min", "expon"]results = []for name in candidates: dist = getattr(stats, name) params = dist.fit(data) log_lik = np.sum(dist.logpdf(data, *params)) k = len(params) aic = 2 * k - 2 * log_lik ks_stat, ks_p = stats.kstest(data, name, args=params) results.append((name, aic, ks_p))results.sort(key=lambda r: r[1]) # lower AIC = better fitfor name, aic, ks_p in results: print(f"{name:12s} AIC={aic:8.1f} KS p={ks_p:.4f}")
Transformations & Order Statistics
Derive the distribution of a function of a random variable, and of sample extremes.
import numpy as npfrom scipy import stats# Transformation of a random variable: if X ~ N(0,1), then Y = X^2 ~ chi-square(1)x = stats.norm.rvs(size=100000, random_state=0)y = x ** 2print(np.mean(y), np.var(y)) # matches chi2(df=1): mean=1, var=2# Log-normal is the exponential transform of a normallog_returns = stats.norm.rvs(loc=0.0005, scale=0.02, size=100000, random_state=1)prices = np.exp(log_returns) # lognormal-distributed# Order statistics: distribution of the sample max of n uniforms is Beta(n, 1)n = 10maxima = stats.uniform.rvs(size=(100000, n)).max(axis=1)theoretical = stats.beta(a=n, b=1)print(f"empirical mean max: {maxima.mean():.4f}, theoretical: {theoretical.mean():.4f}")
Mixture Distributions
Model multimodal data as a weighted combination of component distributions.
import numpy as npfrom sklearn.mixture import GaussianMixture# Simulate bimodal data: two overlapping clustersrng = np.random.default_rng(0)data = np.concatenate([ rng.normal(-2, 0.8, 300), rng.normal(3, 1.2, 200),]).reshape(-1, 1)gmm = GaussianMixture(n_components=2, random_state=0).fit(data)print("component means:", gmm.means_.ravel())print("component weights:", gmm.weights_)print("component variances:", gmm.covariances_.ravel())# Use BIC to select the number of componentsbics = [GaussianMixture(n_components=k, random_state=0).fit(data).bic(data) for k in range(1, 6)]best_k = np.argmin(bics) + 1
Advanced Distribution Reference
Distributions beyond the intro set, useful for heavy tails, rates, and counts.
- Pareto- heavy-tailed power law; models wealth, city sizes, and file sizes (80/20 phenomena)
- Cauchy- symmetric but so heavy-tailed its mean and variance are undefined; a stress test for statistics assuming finite moments
- Weibull- flexible failure-time distribution; shape parameter k controls whether hazard rate increases, decreases, or stays constant
- Gamma- sum of k independent exponential waiting times; generalizes the exponential and chi-square
- Beta- distribution over probabilities in [0,1]; the natural conjugate prior for a Bernoulli/Binomial rate
- Negative binomial- discrete count of failures before r successes; models overdispersed count data where variance > mean (Poisson can't)
- Log-normal- variable whose log is normal; models multiplicative processes like stock prices and income
- Multinomial- generalization of binomial to more than two outcome categories per trial
The Central Limit Theorem means the sampling distribution of a mean approaches normal as sample size grows, regardless of the underlying distribution — this is why t-tests and z-tests work reasonably well even on non-normal data once n is large enough (roughly n ≥ 30).