What Is a Confusion Matrix in Machine Learning
SkillVeris Team
AI Research Team

A confusion matrix is a table that compares a model's predicted labels against the true labels, revealing exactly which classes it confuses.
In this guide, you'll learn:
- For binary classification it has four cells: true positives, true negatives, false positives, and false negatives.
- From these four counts you can compute accuracy, precision, recall, and the F1 score — the core metrics of classification.
- It exposes problems that a single accuracy number hides, especially on imbalanced datasets.
- False positives and false negatives carry different real-world costs, and the matrix makes that trade-off visible.
1What Is a Confusion Matrix?
A confusion matrix is a table that lays out how a classification model's predictions compare to the actual answers. Each row typically represents the true class and each column the predicted class, so every cell counts how many examples of one class were predicted as another. The name comes from its purpose: it shows precisely where the model gets confused.
Unlike a single accuracy score, the matrix breaks performance down class by class. That detail is what makes it the starting point for evaluating any classifier, from spam filters to medical diagnosis models.
2The Four Cells of Binary Classification
For a two-class problem — say, predicting whether an email is spam — the matrix has four cells. Learning these four terms is the key to everything else.
- True Positive (TP): predicted positive and it really is positive (spam correctly flagged).
- True Negative (TN): predicted negative and it really is negative (a normal email correctly allowed through).
- False Positive (FP): predicted positive but it is actually negative (a good email wrongly marked spam) — a 'false alarm'.
- False Negative (FN): predicted negative but it is actually positive (spam that slipped into the inbox) — a 'miss'.
💡A Memory Trick
The second word tells you what the model predicted; the first word tells you whether that prediction was right. 'False Positive' = a positive prediction that was false.
3Metrics You Can Compute From It
The real power of the confusion matrix is that every important classification metric is just arithmetic on its four counts.
- Accuracy = (TP + TN) / total — the share of all predictions that were correct.
- Precision = TP / (TP + FP) — of everything flagged positive, how much really was positive.
- Recall = TP / (TP + FN) — of everything that was actually positive, how much did we catch.
- F1 score = the harmonic mean of precision and recall, balancing the two into one number.
Why Not Just Accuracy?
Accuracy alone can badly mislead. If 99 percent of transactions are legitimate, a model that predicts 'not fraud' every time scores 99 percent accuracy while catching zero fraud. The confusion matrix exposes that failure instantly by showing all the false negatives.
4A Worked Example
Imagine a model that screens 100 emails, of which 20 are truly spam. Suppose it flags 18 emails as spam: 15 are real spam and 3 are innocent. That means 5 spam emails slipped through.
- TP = 15 (spam correctly flagged)
- FP = 3 (good emails wrongly flagged)
- FN = 5 (spam that got through)
- TN = 77 (good emails correctly allowed)
- Precision = 15 / (15 + 3) = 0.83
- Recall = 15 / (15 + 5) = 0.75
🔑Read the Story
This model is fairly precise but misses a quarter of spam. Whether that is acceptable depends entirely on how costly a missed spam email is versus a wrongly blocked one.
5False Positives vs False Negatives
The two kinds of error are rarely equal in the real world, and the confusion matrix forces you to think about which one hurts more. This trade-off should drive how you tune a model.
In cancer screening, a false negative — telling a sick patient they are healthy — is catastrophic, so you tune for high recall even at the cost of more false alarms. In a spam filter, a false positive that deletes an important email may annoy users more than an occasional spam slipping through, so you favor precision. There is no universally correct balance; it depends on consequences.
6Extending to Multiple Classes
When you have more than two classes — say, classifying images as cat, dog, or bird — the confusion matrix becomes a larger square grid, one row and column per class. The principle stays the same.
Correct predictions land on the diagonal, where the true class equals the predicted class. Everything off the diagonal is an error, and reading across a row tells you exactly which other classes a given class gets mistaken for. A cluster of errors between 'cat' and 'dog' cells, for example, signals the model struggles to tell those two apart.
7Building One in Python
In practice you rarely compute a confusion matrix by hand. Libraries like scikit-learn generate it in one line, and add a text report of the derived metrics.
- from sklearn.metrics import confusion_matrix, classification_report
- cm = confusion_matrix(y_true, y_pred) # returns the grid of counts
- print(cm) # rows = actual, columns = predicted
- print(classification_report(y_true, y_pred)) # precision, recall, F1 per class
- import seaborn as sns; sns.heatmap(cm, annot=True) # visualize as a heatmap
8Common Mistakes to Avoid
A confusion matrix is simple to produce but easy to misread. Watch for these traps.
- Confusing the axes — always confirm whether rows or columns represent the true labels before interpreting.
- Judging an imbalanced dataset by accuracy alone instead of reading the matrix.
- Ignoring the off-diagonal cells that reveal which classes get mixed up.
- Optimizing precision and recall in isolation without considering the cost of each error type.
- Evaluating on the training set instead of a held-out test set, which hides overfitting.
9Key Takeaways
Keep these points in mind whenever you evaluate a classifier.
- A confusion matrix compares predicted labels against true labels, class by class.
- Its four binary cells — TP, TN, FP, FN — underpin precision, recall, and F1.
- It reveals failures that a single accuracy figure hides, especially on imbalanced data.
- False positives and false negatives carry different costs; tune your model accordingly.
- The multi-class version puts correct predictions on the diagonal and errors off it.
10Frequently Asked Questions
Q: What does a confusion matrix tell you that accuracy does not? A: It shows the specific types of mistakes a model makes — how many false positives versus false negatives — and which classes get confused with each other. Accuracy compresses all of that into one number that can hide serious problems on imbalanced data.
Q: How do you read a confusion matrix? A: Check which axis holds the true labels and which holds the predictions, then look at the diagonal for correct predictions and the off-diagonal cells for errors. Each off-diagonal cell tells you how often one class was mistaken for another.
Q: Is a confusion matrix only for binary classification? A: No. It extends naturally to any number of classes, becoming a square grid with one row and column per class. The diagonal always represents correct predictions regardless of how many classes you have.
Q: What is the difference between a false positive and a false negative? A: A false positive is a negative example wrongly predicted as positive (a false alarm), while a false negative is a positive example wrongly predicted as negative (a miss). Which is worse depends on the real-world cost of each in your application.
Related Reading
Get The Print Version
Download a PDF of this article for offline reading.
About the Publisher
SkillVeris Team
AI Research Team
Our AI team covers the latest in machine learning, generative AI, and emerging tech — clearly and accurately.
View all postsRelated Posts
Never miss an update
Get the latest tutorials and guides delivered to your inbox.
No spam. Unsubscribe anytime.