Dimensionality Reduction Cheat Sheet
Techniques for reducing feature space while preserving structure, covering PCA, t-SNE, and UMAP with scikit-learn implementation examples.
Principal Component Analysis
Linearly project data onto directions of maximum variance.
from sklearn.decomposition import PCAfrom sklearn.preprocessing import StandardScalerimport numpy as np# Always scale before PCA -- it's sensitive to feature variancescaler = StandardScaler()X_scaled = scaler.fit_transform(X)pca = PCA(n_components=0.95) # keep enough components for 95% varianceX_reduced = pca.fit_transform(X_scaled)print(f"Components kept: {pca.n_components_}")print(f"Explained variance ratio: {pca.explained_variance_ratio_}")print(f"Cumulative variance: {np.cumsum(pca.explained_variance_ratio_)}")
t-SNE & UMAP for Visualization
Non-linear projection into 2D for plotting.
from sklearn.manifold import TSNEimport umapimport matplotlib.pyplot as plt# t-SNE: good for visualizing clusters in 2D/3D, not for feature engineeringtsne = TSNE(n_components=2, perplexity=30, random_state=42)X_tsne = tsne.fit_transform(X_scaled)# UMAP: faster than t-SNE, better preserves global structurereducer = umap.UMAP(n_components=2, n_neighbors=15, min_dist=0.1, random_state=42)X_umap = reducer.fit_transform(X_scaled)plt.scatter(X_umap[:, 0], X_umap[:, 1], c=y, cmap="tab10", s=5)
Dimensionality Reduction Techniques
Common methods and what they're good for.
- PCA- linear technique projecting data onto orthogonal axes of maximum variance
- Explained variance ratio- fraction of total variance captured by each principal component
- t-SNE- non-linear technique optimized for visualizing local cluster structure in 2D/3D
- UMAP- non-linear technique, faster than t-SNE, better preserves both local and global structure
- LDA (Linear Discriminant Analysis)- supervised reduction maximizing class separability
- Autoencoders- neural networks that learn a compressed latent representation via reconstruction
- Feature selection vs extraction- selection keeps a subset of original features, extraction creates new combined features
When to Reach for Each Method
Practical guidance for choosing a technique.
- Speed up training / reduce noise- PCA is fast, linear, and reversible-ish via inverse_transform
- Visualize high-dimensional clusters- t-SNE or UMAP for a 2D/3D scatter plot
- Preserve class separability for a classifier- LDA when labels are available
- Non-linear structure with reconstruction needed- autoencoders for complex, learnable compression
- High curse-of-dimensionality risk- reduce dimensions before distance-based methods like KNN or clustering
Choosing the Number of Components
Use a scree plot and cross-validated reconstruction error instead of picking n_components arbitrarily.
import numpy as npimport matplotlib.pyplot as pltfrom sklearn.decomposition import PCAfrom sklearn.model_selection import KFoldpca_full = PCA().fit(X_scaled)cum_var = np.cumsum(pca_full.explained_variance_ratio_)# Scree plot: look for the 'elbow' where marginal variance gain flattensplt.plot(range(1, len(cum_var) + 1), cum_var, marker="o")plt.axhline(0.95, color="red", linestyle="--")plt.xlabel("Number of components"); plt.ylabel("Cumulative explained variance")# Cross-validated reconstruction error to pick a defensible kdef cv_reconstruction_error(X, k, n_splits=5): errors = [] for train_idx, val_idx in KFold(n_splits).split(X): pca = PCA(n_components=k).fit(X[train_idx]) X_val_rec = pca.inverse_transform(pca.transform(X[val_idx])) errors.append(np.mean((X[val_idx] - X_val_rec) ** 2)) return np.mean(errors)
Kernel PCA & Truncated SVD
Handle non-linear structure and sparse/high-dimensional data where standard PCA falls short.
from sklearn.decomposition import KernelPCA, TruncatedSVD# Kernel PCA: projects into a non-linear feature space via the kernel trickkpca = KernelPCA(n_components=2, kernel="rbf", gamma=0.05, fit_inverse_transform=True)X_kpca = kpca.fit_transform(X_scaled)# Truncated SVD: works directly on sparse matrices (e.g. TF-IDF), unlike PCA# which requires centering (densifying) the datafrom sklearn.feature_extraction.text import TfidfVectorizertfidf = TfidfVectorizer(max_features=5000)X_tfidf = tfidf.fit_transform(corpus)svd = TruncatedSVD(n_components=100, random_state=42)X_svd = svd.fit_transform(X_tfidf) # this is the basis of Latent Semantic Analysis
Supervised & Out-of-Sample UMAP
Use label information to sharpen the embedding and transform new points without refitting.
import umap# Supervised UMAP: uses labels to pull same-class points closer together,# producing a much cleaner embedding for exploratory visualizationreducer = umap.UMAP(n_neighbors=15, min_dist=0.1, random_state=42)embedding = reducer.fit_transform(X_train_scaled, y=y_train)# Unlike t-SNE, a fitted UMAP model supports .transform() on new data --# useful for projecting a validation/production batch into the same spacenew_embedding = reducer.transform(X_new_scaled)# Metric learning: choose a distance metric that matches the datareducer_cosine = umap.UMAP(metric="cosine", n_neighbors=30, random_state=42)
Pitfalls & Key Hyperparameters
The knobs that most often get misused when tuning non-linear reduction methods.
- t-SNE perplexity- roughly the effective number of neighbors considered; try several values (5-50) since results vary significantly with it
- t-SNE cluster sizes/gaps- are not meaningful; the algorithm does not preserve relative cluster size or inter-cluster distance
- UMAP n_neighbors- low values preserve local structure, high values preserve more global structure
- UMAP min_dist- controls how tightly points are packed; lower values create tighter, more separated clusters
- Random state stability- always fix random_state; both t-SNE and UMAP are stochastic and produce different layouts on reruns
- Scaling before reduction- skipping StandardScaler lets high-variance features dominate the distance calculations
- Curse of dimensionality- as dimensions grow, distances between points become less discriminative, which is the core motivation for reduction before KNN/clustering
Manifold Learning Alternatives & Feature Selection
Isomap/LLE for manifold structure, and mutual information as a selection-based alternative to extraction.
from sklearn.manifold import Isomap, LocallyLinearEmbeddingfrom sklearn.feature_selection import mutual_info_classif, SelectKBest# Isomap: preserves geodesic distances along a non-linear manifoldisomap = Isomap(n_neighbors=10, n_components=2)X_isomap = isomap.fit_transform(X_scaled)# Locally Linear Embedding: preserves local linear relationships between neighborslle = LocallyLinearEmbedding(n_neighbors=10, n_components=2, method="standard")X_lle = lle.fit_transform(X_scaled)# Feature SELECTION (keeps original features) as an alternative to extractionmi_scores = mutual_info_classif(X_train, y_train, random_state=42)selector = SelectKBest(mutual_info_classif, k=20).fit(X_train, y_train)X_selected = selector.transform(X_train)
Never use t-SNE or UMAP output as input features for a downstream supervised model, and don't interpret distances between distant clusters in a t-SNE plot as meaningful — both techniques are designed for visualization and preserve local neighborhood structure, not global distances.