Exploratory Data Analysis Cheat Sheet
A systematic workflow for exploring a new dataset, covering summary statistics, distribution plots, correlation analysis, and visualization with pandas and seaborn.
First Look at the Data
Shape, types, and summary statistics.
import pandas as pddf = pd.read_csv("data.csv")df.shape # (rows, columns)df.info() # dtypes, non-null counts, memory usagedf.describe() # count, mean, std, min, quartiles, max for numeric colsdf.describe(include="object") # summary for categorical columnsdf.head()df.isnull().sum()df.nunique() # unique value counts per columndf["category"].value_counts(normalize=True) # class proportions
Visualizing the Data
Distributions, boxplots, and correlations.
import seaborn as snsimport matplotlib.pyplot as plt# Distribution of a numeric variablesns.histplot(df["income"], kde=True, bins=30)# Boxplot to spot outliers/spread by categorysns.boxplot(x="category", y="income", data=df)# Correlation heatmapcorr = df.corr(numeric_only=True)sns.heatmap(corr, annot=True, cmap="coolwarm", center=0)# Pairwise relationships between numeric featuressns.pairplot(df, hue="target", vars=["age", "income", "score"])plt.tight_layout()plt.show()
EDA Checklist
A systematic order for exploring a new dataset.
- Shape & dtypes- confirm row/column counts and correct data types before anything else
- Summary statistics- mean, median, std, min/max reveal scale and possible errors
- Missing data pattern- check whether missingness is random or correlated with other features
- Distribution shape- check skewness, multi-modality, and outliers with histograms/KDE
- Correlation analysis- look for multicollinearity and target-feature relationships
- Class balance- check the target variable distribution for classification tasks
- Categorical cardinality- count unique values per categorical column to plan encoding
- Bivariate analysis- explore relationships between pairs of features (scatter, boxplot, groupby)
Useful Pandas Methods
Quick lookups for common EDA operations.
- df.groupby(col).agg()- compute grouped summary statistics
- df.corr()- pairwise correlation matrix of numeric columns
- df.value_counts()- frequency count of unique values in a Series
- df.sample(n)- random subset of rows for a quick sanity check
- df.select_dtypes()- filter columns by data type (e.g. include=['number'])
- pd.crosstab()- cross-tabulation frequency table between two categorical columns
Statistical Tests for Feature-Target Relationships
Quantify whether an observed relationship in EDA plots is likely real or noise.
from scipy import stats# Numeric feature vs numeric target: Pearson/Spearman correlation with p-valuer, p = stats.pearsonr(df["income"], df["score"])rho, p_s = stats.spearmanr(df["income"], df["score"]) # robust to non-linearity# Numeric feature vs categorical target: one-way ANOVAgroups = [g["income"].values for _, g in df.groupby("target")]f_stat, p_anova = stats.f_oneway(*groups)# Categorical feature vs categorical target: chi-square test of independencecontingency = pd.crosstab(df["category"], df["target"])chi2, p_chi2, dof, expected = stats.chi2_contingency(contingency)print(f"pearson p={p:.4f} anova p={p_anova:.4f} chi2 p={p_chi2:.4f}")
Skewness, Kurtosis & Normality Diagnostics
Quantify distribution shape instead of eyeballing histograms alone.
from scipy.stats import skew, kurtosis, shapiro, normaltestfor col in ["income", "score", "age"]: s = df[col].dropna() print(col, f"skew={skew(s):.2f}", # >1 or <-1 = highly skewed f"kurtosis={kurtosis(s):.2f}", # excess kurtosis; >0 = heavy tails )# Shapiro-Wilk normality test (best for n < 5000)stat, p = shapiro(df["income"].sample(min(len(df), 5000), random_state=42))print("normal" if p > 0.05 else "not normal", f"(p={p:.4f})")# D'Agostino's K^2 test, works well for larger samplesstat2, p2 = normaltest(df["income"].dropna())
Multi-Index GroupBy & Pivot Tables
Slice a dataset along several dimensions at once during exploratory passes.
# Multiple aggregations per column, grouped by multiple keyssummary = df.groupby(["region", "category"]).agg( avg_income=("income", "mean"), median_score=("score", "median"), n=("income", "size"), std_income=("income", "std"),)# Pivot table with margins for row/column totalspivot = pd.pivot_table( df, values="income", index="region", columns="category", aggfunc="mean", margins=True, margins_name="Overall",)# Rank within group -- useful for spotting top performers per segmentdf["income_rank_in_region"] = df.groupby("region")["income"].rank(ascending=False)
Which Statistical Test to Use
Quick reference for matching a test to the variable types you're comparing.
- Numeric vs numeric- Pearson correlation (linear) or Spearman correlation (monotonic, robust to outliers)
- Numeric vs binary categorical- independent t-test (or Mann-Whitney U if non-normal)
- Numeric vs multi-class categorical- one-way ANOVA (or Kruskal-Wallis if non-normal)
- Categorical vs categorical- chi-square test of independence on a contingency table
- Distribution shape check- Shapiro-Wilk or D'Agostino K^2 test for normality
- Two-sample distribution comparison- Kolmogorov-Smirnov (KS) test, e.g. train vs test drift
- Multicollinearity check- Variance Inflation Factor (VIF); VIF > 10 signals a redundant feature
Time Series EDA
Rolling statistics and decomposition to explore trend, seasonality, and autocorrelation.
from statsmodels.tsa.seasonal import seasonal_decomposefrom statsmodels.graphics.tsaplots import plot_acf, plot_pacfts = df.set_index("date")["sales"].asfreq("D").interpolate()# Rolling mean/std to visualize trend and changing variancerolling_mean = ts.rolling(window=30).mean()rolling_std = ts.rolling(window=30).std()# Decompose into trend, seasonal, and residual componentsdecomposition = seasonal_decompose(ts, model="additive", period=365)decomposition.plot()# Autocorrelation / partial autocorrelation -- guides ARIMA (p, q) orderplot_acf(ts.dropna(), lags=60)plot_pacf(ts.dropna(), lags=60)
Always plot the target variable's distribution first — a skewed regression target (e.g. housing prices) often benefits from a log transform, and a heavily imbalanced classification target changes which metrics and resampling strategies you should use downstream.