How to Handle Missing Data in a Dataset
SkillVeris Team
Data Science Team

Handle missing data by first diagnosing why it is missing, then choosing between deletion and imputation based on how much is missing and the mechanism behind it.
In this guide, you'll learn:
- The three missingness mechanisms — MCAR, MAR, and MNAR — determine which methods are safe to use.
- Deletion is simple but discards information; imputation preserves rows but adds assumptions.
- Simple imputation fills gaps with the mean, median, or mode; advanced methods model the missing values.
- Always quantify missingness per column before acting — isna().sum() is the first step.
1How to Handle Missing Data
To handle missing data well, first understand why values are missing, then choose between deleting the affected data and imputing (filling in) the gaps. The right choice depends on how much data is missing, which columns are affected, and the mechanism causing the absence. There is no single correct method — only trade-offs you should make deliberately.
Missing values are one of the most common problems in real datasets, and how you handle them directly shapes your results. Delete too aggressively and you throw away information; impute carelessly and you bias the data. The goal is to make an informed decision rather than reaching for a reflexive default.
2Diagnose the Missingness First
Before fixing anything, measure it. Count how many values are missing in each column and what fraction of the data that represents. A column that is 2% missing invites a different strategy than one that is 60% missing, which may be worth dropping entirely. Visualizing the pattern also reveals whether values go missing together, which is a clue to the underlying cause.
- df.isna().sum() # count missing per column
- df.isna().mean().round(3) # fraction missing per column
- df.isna().sum().sum() # total missing cells
- df[df['income'].isna()] # inspect rows missing income
💡Percentages Guide the Choice
Convert counts to percentages with isna().mean(). A column that is mostly empty is often better dropped than imputed, while a column missing a few percent is a good candidate for imputation.
3Understand Why Data Is Missing
Statisticians describe three missingness mechanisms, and knowing which you face keeps your fix honest. The mechanism determines whether simply deleting or mean-filling will bias your results or leave them intact.
- MCAR (Missing Completely At Random): the gap is unrelated to anything — safest to delete.
- MAR (Missing At Random): missingness depends on other observed columns — imputation can work.
- MNAR (Missing Not At Random): missingness depends on the missing value itself — hardest, and easy to bias.
- Example of MNAR: high earners declining to report income skews any naive fill.
4Deletion Methods
Deletion removes data containing missing values. Listwise deletion drops any row with a missing value; column deletion drops a whole feature that is mostly empty. Deletion is simple and introduces no invented values, but it discards information and can bias results if the missingness is not completely at random. It is reasonable when only a small fraction of rows are affected.
- df.dropna() # drop rows with any missing value
- df.dropna(subset=['age', 'income']) # only where these are missing
- df.dropna(axis=1, thresh=len(df)*0.5) # drop columns >50% empty
- df.dropna(how='all') # drop rows that are entirely empty
⚠️Deletion Can Bias
If the rows you delete differ systematically from the rest — for example, low-income users skip the income field — dropping them skews your sample. Deletion is only safe when missingness is essentially random.
5Imputation Methods
Imputation fills missing values with estimates so you keep every row. Simple imputation uses a summary statistic — mean or median for numbers, mode for categories. More advanced approaches predict the missing value from other columns, using techniques like K-nearest-neighbors or iterative modeling. Simple methods are fast and often good enough; model-based methods are more accurate but add complexity and their own assumptions.
Simple Imputation in Pandas
For many datasets, median and mode filling handle the majority of cases cleanly.
df['age'].fillna(df['age'].median()) # robust to outliers
df['income'].fillna(df['income'].mean())
df['city'].fillna(df['city'].mode()[0]) # most common category
df['score'].fillna(method='ffill') # carry last value forward (time series)Model-Based Imputation
Scikit-learn offers imputers that estimate missing values from the other features rather than a single global statistic.
from sklearn.impute import KNNImputer
imputer = KNNImputer(n_neighbors=5)
X_filled = imputer.fit_transform(X_train) # fit on training data only6Preserve the Signal in Missingness
Sometimes the fact that a value is missing is itself informative. A customer who skips the 'annual income' field may differ systematically from one who fills it in. Before you impute away the gaps, consider adding a binary indicator column that records where the value was originally missing. That way the model can still learn from the pattern of absence, even after you fill in the numbers.
- df['income_missing'] = df['income'].isna().astype(int) # flag before filling
- df['income'] = df['income'].fillna(df['income'].median())
7Common Mistakes to Avoid
These missteps quietly corrupt datasets and downstream models.
- Filling gaps before understanding why the data is missing.
- Using the mean on skewed columns where the median is far more representative.
- Imputing using statistics from the whole dataset, leaking test data into training.
- Dropping rows with any missing value and discarding most of the data.
- Ignoring that missingness itself can be a meaningful signal worth flagging.
8Key Takeaways
Handle missing data with these principles in mind.
- Diagnose first: count and visualize missingness before acting.
- The mechanism (MCAR, MAR, MNAR) determines which methods are safe.
- Deletion is simple but discards data and can bias results.
- Impute with median or mode simply, or model-based methods for accuracy.
- Fit imputers on training data only, and consider a 'was missing' indicator.
9Frequently Asked Questions
Q: Should I delete or impute missing data? A: It depends on how much is missing and why. Delete when only a small fraction of rows are affected and the missingness is essentially random. Impute when dropping would discard too much data, choosing a method that fits the column and mechanism.
Q: What is the best imputation method? A: There is no universal best. Median works well for skewed numeric columns, mode for categories, and model-based methods like KNN when relationships between columns are strong. The right choice balances accuracy against complexity for your data.
Q: What do MCAR, MAR, and MNAR mean? A: They describe why data is missing. MCAR means the gap is unrelated to anything, MAR means it depends on other observed columns, and MNAR means it depends on the missing value itself. MNAR is the hardest to handle without bias.
Q: Why avoid using the mean to fill skewed data? A: The mean is pulled toward extreme values, so on skewed columns it can misrepresent the typical value and distort the distribution. The median is robust to outliers and usually a safer fill for skewed numeric data.
Related Reading
Get The Print Version
Download a PDF of this article for offline reading.
About the Publisher
SkillVeris Team
Data Science Team
Our data team shares real-world analytics, ML, and SQL insights grounded in industry practice.
View all postsRelated Posts
Never miss an update
Get the latest tutorials and guides delivered to your inbox.
No spam. Unsubscribe anytime.