How to Perform Exploratory Data Analysis in Python
SkillVeris Team
Data Science Team

Exploratory data analysis (EDA) is the process of summarizing and visualizing a dataset to understand its structure, spot problems, and form hypotheses before modeling.
In this guide, you'll learn:
- A good EDA answers: what does each variable look like, how do variables relate, and what is missing or anomalous?
- Start with univariate analysis (one variable), then bivariate (two variables), then multivariate patterns.
- Pandas plus matplotlib or seaborn cover the whole workflow — describe, then plot distributions and relationships.
- Visualization reveals patterns that summary statistics hide, like skew, clusters, and outliers.
1What Is Exploratory Data Analysis?
Exploratory data analysis (EDA) is the practice of summarizing and visualizing a dataset to understand its main characteristics before you build models or draw conclusions. It answers three core questions: what does each variable look like on its own, how do variables relate to each other, and what is missing, wrong, or surprising?
The term was popularized by statistician John Tukey, who argued that looking at data openly, without a fixed hypothesis, reveals structure you would otherwise miss. EDA is where you build intuition and catch problems early, so the modeling that follows rests on solid ground.
2Getting a First Look
Every EDA starts with orientation: load the data and understand its shape, types, and summary statistics. These first commands tell you how big the dataset is, which columns are numeric or categorical, and where values are missing. Do this before any plotting so you know what you are working with.
- import pandas as pd
- df = pd.read_csv('data.csv')
- df.shape # rows, columns
- df.info() # types and non-null counts
- df.describe() # numeric summary
- df.describe(include='object') # categorical summary
- df.isna().sum() # missing values per column
3Univariate Analysis: One Variable at a Time
Univariate analysis examines each variable on its own to understand its distribution. For numeric columns, a histogram shows the shape — is it symmetric, skewed, or bimodal — and a box plot reveals spread and outliers. For categorical columns, value_counts and a bar chart show how frequent each category is. This step tells you what normal looks like for every field.
- df['age'].hist(bins=30) # distribution shape
- df['age'].plot(kind='box') # spread and outliers
- df['category'].value_counts() # category frequencies
- df['category'].value_counts().plot(kind='bar')
💡Plot the Distribution
Summary statistics can hide a lot. Two columns with the same mean can look completely different — always plot the histogram to see the real shape.
4Bivariate Analysis: Relationships Between Variables
Bivariate analysis looks at how two variables relate. A scatter plot reveals whether two numeric variables move together. A correlation matrix quantifies those linear relationships across all numeric columns at once. Grouping a numeric variable by a category and comparing distributions shows how categories differ. These relationships are exactly what predictive models will try to exploit.
- df.plot.scatter(x='height', y='weight') # numeric vs numeric
- df.corr(numeric_only=True) # correlation matrix
- df.groupby('gender')['income'].mean() # numeric by category
- df.boxplot(column='income', by='gender') # compare distributions
⚠️Correlation Is Not Causation
A strong correlation shows two variables move together, not that one causes the other. A hidden third variable often drives both. Treat correlations as leads, not conclusions.
5Choosing Visualization Tools
Two libraries cover almost all EDA plotting in Python, and they complement each other. Pandas has quick built-in plotting for fast exploration. Seaborn, built on matplotlib, produces attractive statistical charts with far less code and handles grouping elegantly. Reach for whichever gets you to insight fastest.
Seaborn for Statistical Plots
Seaborn shines for relationships and distributions across categories in a single call.
import seaborn as sns
sns.histplot(df, x='age', hue='gender') # grouped distribution
sns.heatmap(df.corr(numeric_only=True), annot=True) # correlations
sns.pairplot(df[['age', 'income', 'score']]) # all pairs at once6EDA Is Iterative
EDA is not a checklist you complete once; it is a loop. Each chart raises a new question. A skewed distribution prompts you to check for outliers. An odd correlation prompts you to split the data by a category. You keep cycling — summarize, visualize, question, dig deeper — until the dataset holds no more surprises and you can state clearly what it contains and how its parts relate.
7Common Mistakes to Avoid
Avoid these traps that produce a shallow or misleading analysis.
- Jumping straight to modeling without understanding the data first.
- Relying on summary statistics alone and never plotting the distributions.
- Reading correlation as causation and drawing unfounded conclusions.
- Ignoring missing values and outliers that distort every chart.
- Treating EDA as one-and-done instead of iterating on each new question.
8Key Takeaways
Keep the EDA workflow tight with these principles.
- EDA summarizes and visualizes data to understand it before modeling.
- Work in order: univariate, then bivariate, then multivariate.
- Plot distributions — statistics alone hide skew, clusters, and outliers.
- Use pandas for quick plots and seaborn for statistical charts.
- Iterate: each finding raises the next question to explore.
9Frequently Asked Questions
Q: What is the goal of exploratory data analysis? A: To understand a dataset's structure, quality, and relationships before modeling. Good EDA reveals distributions, correlations, missing values, and anomalies, so you catch problems early and form well-grounded hypotheses.
Q: What tools do I need for EDA in Python? A: Pandas for loading and summarizing data, plus matplotlib or seaborn for visualization, cover the entire workflow. Jupyter or a similar notebook environment makes the iterative loop of code and charts smooth.
Q: What is the difference between univariate and bivariate analysis? A: Univariate analysis examines one variable at a time to understand its distribution. Bivariate analysis examines two variables together to see how they relate, using scatter plots, correlations, and grouped comparisons.
Q: How long should EDA take? A: It varies with dataset size and complexity, but plan for a meaningful chunk of the project. EDA is iterative, and rushing it usually leads to modeling on misunderstood data, which costs far more time later.
Related Reading
Get The Print Version
Download a PDF of this article for offline reading.
About the Publisher
SkillVeris Team
Data Science Team
Our data team shares real-world analytics, ML, and SQL insights grounded in industry practice.
View all postsRelated Posts
Never miss an update
Get the latest tutorials and guides delivered to your inbox.
No spam. Unsubscribe anytime.