100% Free Forever
AI-Powered Learning
Industry Expert Content
Certificates & Badges
Learn At Your Own Pace
HomeBlogHow to Handle Missing Data in a Dataset
Data Science

How to Handle Missing Data in a Dataset

SV

SkillVeris Team

Data Science Team

Nov 24, 2025 10 min read
Share:
How to Handle Missing Data in a Dataset
Key Takeaway

Handle missing data by first diagnosing why it is missing, then choosing between deletion and imputation based on how much is missing and the mechanism behind it.

In this guide, you'll learn:

  • The three missingness mechanisms — MCAR, MAR, and MNAR — determine which methods are safe to use.
  • Deletion is simple but discards information; imputation preserves rows but adds assumptions.
  • Simple imputation fills gaps with the mean, median, or mode; advanced methods model the missing values.
  • Always quantify missingness per column before acting — isna().sum() is the first step.

1How to Handle Missing Data

To handle missing data well, first understand why values are missing, then choose between deleting the affected data and imputing (filling in) the gaps. The right choice depends on how much data is missing, which columns are affected, and the mechanism causing the absence. There is no single correct method — only trade-offs you should make deliberately.

Missing values are one of the most common problems in real datasets, and how you handle them directly shapes your results. Delete too aggressively and you throw away information; impute carelessly and you bias the data. The goal is to make an informed decision rather than reaching for a reflexive default.

2Diagnose the Missingness First

Before fixing anything, measure it. Count how many values are missing in each column and what fraction of the data that represents. A column that is 2% missing invites a different strategy than one that is 60% missing, which may be worth dropping entirely. Visualizing the pattern also reveals whether values go missing together, which is a clue to the underlying cause.

  • df.isna().sum() # count missing per column
  • df.isna().mean().round(3) # fraction missing per column
  • df.isna().sum().sum() # total missing cells
  • df[df['income'].isna()] # inspect rows missing income

💡Percentages Guide the Choice

Convert counts to percentages with isna().mean(). A column that is mostly empty is often better dropped than imputed, while a column missing a few percent is a good candidate for imputation.

3Understand Why Data Is Missing

Statisticians describe three missingness mechanisms, and knowing which you face keeps your fix honest. The mechanism determines whether simply deleting or mean-filling will bias your results or leave them intact.

  • MCAR (Missing Completely At Random): the gap is unrelated to anything — safest to delete.
  • MAR (Missing At Random): missingness depends on other observed columns — imputation can work.
  • MNAR (Missing Not At Random): missingness depends on the missing value itself — hardest, and easy to bias.
  • Example of MNAR: high earners declining to report income skews any naive fill.

4Deletion Methods

Deletion removes data containing missing values. Listwise deletion drops any row with a missing value; column deletion drops a whole feature that is mostly empty. Deletion is simple and introduces no invented values, but it discards information and can bias results if the missingness is not completely at random. It is reasonable when only a small fraction of rows are affected.

  • df.dropna() # drop rows with any missing value
  • df.dropna(subset=['age', 'income']) # only where these are missing
  • df.dropna(axis=1, thresh=len(df)*0.5) # drop columns >50% empty
  • df.dropna(how='all') # drop rows that are entirely empty

⚠️Deletion Can Bias

If the rows you delete differ systematically from the rest — for example, low-income users skip the income field — dropping them skews your sample. Deletion is only safe when missingness is essentially random.

5Imputation Methods

Imputation fills missing values with estimates so you keep every row. Simple imputation uses a summary statistic — mean or median for numbers, mode for categories. More advanced approaches predict the missing value from other columns, using techniques like K-nearest-neighbors or iterative modeling. Simple methods are fast and often good enough; model-based methods are more accurate but add complexity and their own assumptions.

Simple Imputation in Pandas

For many datasets, median and mode filling handle the majority of cases cleanly.

code
df['age'].fillna(df['age'].median())    # robust to outliers
df['income'].fillna(df['income'].mean())
df['city'].fillna(df['city'].mode()[0])  # most common category
df['score'].fillna(method='ffill')       # carry last value forward (time series)

Model-Based Imputation

Scikit-learn offers imputers that estimate missing values from the other features rather than a single global statistic.

code
from sklearn.impute import KNNImputer
imputer = KNNImputer(n_neighbors=5)
X_filled = imputer.fit_transform(X_train)  # fit on training data only

6Preserve the Signal in Missingness

Sometimes the fact that a value is missing is itself informative. A customer who skips the 'annual income' field may differ systematically from one who fills it in. Before you impute away the gaps, consider adding a binary indicator column that records where the value was originally missing. That way the model can still learn from the pattern of absence, even after you fill in the numbers.

  • df['income_missing'] = df['income'].isna().astype(int) # flag before filling
  • df['income'] = df['income'].fillna(df['income'].median())

7Common Mistakes to Avoid

These missteps quietly corrupt datasets and downstream models.

  • Filling gaps before understanding why the data is missing.
  • Using the mean on skewed columns where the median is far more representative.
  • Imputing using statistics from the whole dataset, leaking test data into training.
  • Dropping rows with any missing value and discarding most of the data.
  • Ignoring that missingness itself can be a meaningful signal worth flagging.

8Key Takeaways

Handle missing data with these principles in mind.

  • Diagnose first: count and visualize missingness before acting.
  • The mechanism (MCAR, MAR, MNAR) determines which methods are safe.
  • Deletion is simple but discards data and can bias results.
  • Impute with median or mode simply, or model-based methods for accuracy.
  • Fit imputers on training data only, and consider a 'was missing' indicator.

9Frequently Asked Questions

Q: Should I delete or impute missing data? A: It depends on how much is missing and why. Delete when only a small fraction of rows are affected and the missingness is essentially random. Impute when dropping would discard too much data, choosing a method that fits the column and mechanism.

Q: What is the best imputation method? A: There is no universal best. Median works well for skewed numeric columns, mode for categories, and model-based methods like KNN when relationships between columns are strong. The right choice balances accuracy against complexity for your data.

Q: What do MCAR, MAR, and MNAR mean? A: They describe why data is missing. MCAR means the gap is unrelated to anything, MAR means it depends on other observed columns, and MNAR means it depends on the missing value itself. MNAR is the hardest to handle without bias.

Q: Why avoid using the mean to fill skewed data? A: The mean is pulled toward extreme values, so on skewed columns it can misrepresent the typical value and distort the distribution. The median is robust to outliers and usually a safer fill for skewed numeric data.

📄

Get The Print Version

Download a PDF of this article for offline reading.

About the Publisher

SV

SkillVeris Team

Data Science Team

Our data team shares real-world analytics, ML, and SQL insights grounded in industry practice.

View all posts

Never miss an update

Get the latest tutorials and guides delivered to your inbox.

No spam. Unsubscribe anytime.

Frequently Asked Questions

21 categories · pick one to explore

Does SkillVeris have a tech blog, and what does it cover?
Yes, the SkillVeris blog has over 500 articles covering AI and machine learning, programming, web development, DevOps, cloud, security, databases and career guidance. Articles are practical and answer-first, and many use the Learn Through Hobbies approach, teaching technical concepts through cricket, music, gaming or cooking analogies. Everything is free to read.
What is the SkillVeris tech glossary and how big is it?
The SkillVeris glossary is a free reference of roughly 2,000-plus technology terms, each with a clear plain-language definition. It spans AI, programming, web, DevOps, cloud, security and database vocabulary, so whenever a lesson, article or job description uses jargon you do not recognise, the glossary gives you a fast, reliable answer.
Are the developer cheat sheets on SkillVeris free to download?
The cheat sheets are completely free to use, like everything else on SkillVeris. Each sheet condenses a language or tool into its essential syntax, commands and patterns for quick reference while coding. They are designed for rapid lookup during real work, complementing the deeper explanations found in study notes and courses.
Which programming references and cheat sheets are available?
Cheat sheets cover the platform's main domains, including programming languages, AI and ML tooling, web development, DevOps, cloud, security and databases, matching the topics of the 37 live courses. Each sheet lists related reading links and hashtags, so you can jump from a quick reference into fuller study notes or blog articles.
How do I find the meaning of a technical term quickly?
Search the SkillVeris glossary, which holds around 2,000-plus terms with concise, plain-language definitions. Each entry gets to the point in its first sentence, then links to related reading like blog posts or study notes for deeper context. It is faster and more consistent than sifting through scattered search results.
Is the SkillVeris blog good for beginners learning to code?
Yes, many blog articles are written specifically for beginners, and the Learn Through Hobbies style makes them unusually approachable: you might learn Python concepts through cricket or understand APIs through cooking. With 500-plus articles across skill levels, beginners can start with fundamentals and keep reading as they advance, entirely free.
Can cheat sheets replace full courses for learning a language?
No, cheat sheets are references, not teaching tools; they assume you already understand the concepts and just need syntax or commands fast. To actually learn a language, take a structured SkillVeris course with its 24–40 lessons and assessments, then keep the cheat sheet beside you while practising in Code Lab.
How often are new blog articles published on SkillVeris?
The blog grows regularly and already exceeds 500 articles, with new posts added as courses launch and technologies evolve. Topics track the platform's catalogue across AI, programming, web development, DevOps, cloud and security, so checking the Blog section periodically surfaces fresh tutorials, explainers and career-focused pieces, all free to read.
Does the glossary cover AI and machine learning terms?
Yes, AI and machine learning vocabulary is a major part of the roughly 2,000-plus term glossary, covering everything from foundational terms to modern concepts around LLMs, RAG and MLOps. Definitions are plain-language and answer-first, which helps when dense AI papers or course lessons throw unfamiliar jargon at you.
Are there cheat sheets for interview preparation?
Cheat sheets work well as interview-day refreshers because they compress syntax, commands and key concepts into scannable references. For dedicated preparation, combine them with the SkillVeris interview questions feature, which includes readiness scoring, plus study notes for depth. Reviewing a relevant cheat sheet just before an interview steadies recall under pressure.
Can I read the tech blog without signing up?
Yes, the blog is freely readable, and SkillVeris never charges for content. All 500-plus articles are open, covering tutorials, concept explainers and career advice. Creating a free account adds value elsewhere on the platform, like course progress tracking and certificates, but reading the blog requires no commitment at all.
How is the SkillVeris glossary different from Wikipedia?
The glossary is purpose-built for learners: definitions are short, plain-language and answer-first, sized for a quick lookup mid-lesson rather than a deep encyclopedic read. Entries also cross-link to related SkillVeris study notes, blog posts and courses, so a definition becomes a doorway into structured learning instead of a dead end.
Do blog articles use the Learn Through Hobbies method?
Many blog articles teach technical topics through hobby analogies, a hallmark of the SkillVeris blog, so you will find articles explaining programming through cricket, machine learning through music, or system design through cooking. The analogy is the teaching device; the article still delivers the real technical concept underneath.
Where can I find quick programming references while coding?
Open the SkillVeris cheat sheets, which are built exactly for that moment: compact, scannable references for syntax, commands and common patterns across languages and tools. Keep the relevant sheet in a browser tab while you work in Code Lab or your own editor, and dip into the glossary for terminology.
Is there a glossary entry for terms I meet in job descriptions?
Very likely yes, with roughly 2,000-plus terms across AI, programming, web, DevOps, cloud, security and databases, the glossary covers most jargon that appears in tech job descriptions. Decoding a listing this way helps you judge role fit honestly and prepares you to discuss those terms in interviews.
Are the blog articles written for the Indian tech audience?
The blog serves Indian learners plus a worldwide audience. Content stays globally relevant while acknowledging realities that matter in India, such as free access being essential for students and freshers, and career guidance that connects naturally to the SkillVeris jobs portal, which aggregates roles across India, UK, USA, Germany and Remote.
Can I suggest a topic for the blog or glossary?
SkillVeris content grows in response to what learners need, so feedback is welcome through the platform's support channels. If a term is missing from the glossary or a topic deserves an article, telling the team helps prioritise it. Meanwhile, the AI Mentor can answer the question immediately, 24/7, at any depth.
Do cheat sheets and glossary entries link to deeper learning?
Yes, every cheat sheet and glossary entry carries related reading links into study notes, blog articles and courses, plus concept hashtags for discovering similar content. This cross-linking means a thirty-second lookup can smoothly become a structured learning session whenever you decide you want more than a quick answer.
What makes SkillVeris programming references trustworthy?
The references are written to strict internal quality standards, kept consistent with the platform's 37 live courses, and never padded with invented statistics or hype. Definitions and cheat sheets are reviewed against the same content contracts that govern courses, and the answer-first style makes any inaccuracy easy to spot and correct.
How do the blog, glossary and cheat sheets fit into my learning routine?
Use them as satellites around your main course: read blog articles for context and motivation, hit the glossary the instant jargon appears, and keep cheat sheets open while coding. Together with study notes, Code Lab and the 24/7 AI Mentor, they turn passive reading into a complete, free learning system.

What Learners Say

Real journeys from the SkillVeris community — swipe for more.

SkillVeris taught me Python through Cricket. Now I’m building real projects and feeling confident!
Arjun S. · B.Tech Student
The best platform for hobby-based learning. Concepts finally stick.
Priya R. · Data Analyst
I went from zero coding to a portfolio of projects — all by learning through my love for gaming. Landed my first internship!
Kabir M. · CS Undergraduate
Trending Topics50 popular tags — tap to explore
Trending CoursesAll 37 free courses — tap to browse