Confusion Matrix & Metrics Cheat Sheet
How to read a confusion matrix and compute precision, recall, F1, and accuracy for binary and multi-class classification with scikit-learn.
Confusion Matrix & Report
Build and visualize a confusion matrix.
from sklearn.metrics import confusion_matrix, classification_report, ConfusionMatrixDisplayimport matplotlib.pyplot as plty_true = [1, 0, 1, 1, 0, 1, 0, 0]y_pred = [1, 0, 0, 1, 0, 1, 1, 0]cm = confusion_matrix(y_true, y_pred)print(cm)# [[TN FP]# [FN TP]]print(classification_report(y_true, y_pred, target_names=["neg", "pos"]))ConfusionMatrixDisplay(cm, display_labels=["neg", "pos"]).plot()plt.show()
Metric Formulas
How each metric is derived from TP/TN/FP/FN.
# TP = true positive, TN = true negative# FP = false positive (Type I error), FN = false negative (Type II error)accuracy = (TP + TN) / (TP + TN + FP + FN)precision = TP / (TP + FP) # of predicted positives, how many correctrecall = TP / (TP + FN) # of actual positives, how many found (sensitivity)specificity = TN / (TN + FP) # of actual negatives, how many foundf1_score = 2 * (precision * recall) / (precision + recall)
Metric Definitions
Core classification evaluation terms.
- True Positive (TP)- model correctly predicts the positive class
- True Negative (TN)- model correctly predicts the negative class
- False Positive (FP)- model predicts positive but actual is negative (Type I error)
- False Negative (FN)- model predicts negative but actual is positive (Type II error)
- Precision- TP / (TP + FP); how trustworthy positive predictions are
- Recall (Sensitivity)- TP / (TP + FN); how many actual positives were caught
- F1 Score- harmonic mean of precision and recall; good for imbalanced classes
- Accuracy- overall fraction correct; can be misleading on imbalanced datasets
Which Metric to Prioritize
Choosing metrics based on the cost of errors.
- High cost of false positives- optimize for precision (e.g. spam filtering)
- High cost of false negatives- optimize for recall (e.g. cancer screening)
- Balanced classes, balanced costs- accuracy is a reasonable single-number summary
- Imbalanced classes- prefer F1, precision-recall curves, or balanced accuracy over raw accuracy
- Multi-class problems- use macro/micro/weighted averages of precision, recall, and F1
Multi-Class Averaging Strategies
Compute macro, micro, and weighted precision/recall/F1 for multi-class problems.
from sklearn.metrics import precision_recall_fscore_support, confusion_matriximport numpy as npy_true = [0, 1, 2, 2, 1, 0, 2, 1, 0, 2]y_pred = [0, 2, 2, 2, 1, 0, 1, 1, 0, 2]cm = confusion_matrix(y_true, y_pred)print(cm) # rows = actual, cols = predicted# macro: unweighted mean across classes (treats rare classes equally)macro = precision_recall_fscore_support(y_true, y_pred, average="macro")# weighted: mean weighted by class support (accounts for imbalance)weighted = precision_recall_fscore_support(y_true, y_pred, average="weighted")# micro: aggregate TP/FP/FN globally first, then compute (== accuracy for single-label multi-class)micro = precision_recall_fscore_support(y_true, y_pred, average="micro")print(f"macro P/R/F1: {macro[:3]}")print(f"weighted P/R/F1: {weighted[:3]}")print(f"micro P/R/F1: {micro[:3]}")
Matthews Correlation Coefficient & Cohen's Kappa
Single-number metrics that are robust to class imbalance, unlike accuracy.
from sklearn.metrics import matthews_corrcoef, cohen_kappa_score, balanced_accuracy_scorey_true = [1, 0, 1, 1, 0, 1, 0, 0, 1, 0]y_pred = [1, 0, 0, 1, 0, 1, 1, 0, 1, 0]# MCC: correlation between predicted and actual (-1..1), works even on# heavily imbalanced binary/multi-class data where F1 can misleadmcc = matthews_corrcoef(y_true, y_pred)# Cohen's kappa: agreement between prediction and truth, corrected for# the agreement expected by chance alone (useful for inter-rater comparison too)kappa = cohen_kappa_score(y_true, y_pred)# balanced accuracy: average of recall obtained on each classbal_acc = balanced_accuracy_score(y_true, y_pred)print(f"MCC: {mcc:.3f}, Kappa: {kappa:.3f}, Balanced Accuracy: {bal_acc:.3f}")
Cost-Sensitive Threshold Selection
Pick a decision threshold that minimizes business cost instead of the default 0.5.
import numpy as npfrom sklearn.metrics import confusion_matrixy_scores = model.predict_proba(X_test)[:, 1]cost_fp, cost_fn = 5, 25 # false negative is 5x more costly herebest_threshold, best_cost = 0.5, np.inffor t in np.arange(0.05, 0.96, 0.01): y_pred = (y_scores >= t).astype(int) tn, fp, fn, tp = confusion_matrix(y_test, y_pred).ravel() total_cost = fp * cost_fp + fn * cost_fn if total_cost < best_cost: best_cost, best_threshold = total_cost, tprint(f"Optimal threshold: {best_threshold:.2f}, expected cost: {best_cost}")
Per-Label Confusion Matrices
Get one confusion matrix per label for multi-label classification.
from sklearn.metrics import multilabel_confusion_matrix, classification_report# each row is a sample that can belong to multiple labels at oncey_true = [[1, 0, 1], [0, 1, 0], [1, 1, 0]]y_pred = [[1, 0, 0], [0, 1, 1], [1, 0, 0]]mcm = multilabel_confusion_matrix(y_true, y_pred)# mcm[i] is a 2x2 [[TN, FP], [FN, TP]] matrix for label ifor i, m in enumerate(mcm): print(f"label {i}:\n{m}")print(classification_report(y_true, y_pred, target_names=["lbl0", "lbl1", "lbl2"]))
Advanced Metrics Glossary
Less common but powerful single-number classification metrics.
- Matthews Correlation Coefficient (MCC)- ranges -1 to 1; considered the most reliable single score on imbalanced binary data because it uses all four confusion matrix cells
- Cohen's Kappa- agreement score corrected for chance; also used to compare two human labelers or a model vs. a labeler
- Balanced Accuracy- average of per-class recall; equals accuracy on balanced data but doesn't reward majority-class bias on imbalanced data
- Youden's J statistic- sensitivity + specificity - 1; a fast way to pick a threshold that balances both error types
- G-Mean- geometric mean of sensitivity and specificity; sensitive to poor performance on either class
- Fowlkes-Mallows Index- geometric mean of precision and recall; common in clustering evaluation, applicable to classification too
- Micro vs. macro F1- micro is dominated by frequent classes, macro treats every class equally regardless of size
On imbalanced datasets, accuracy can look great while the model just predicts the majority class every time — always check precision, recall, and F1 for the minority class, or look at the confusion matrix directly.