Feature Engineering Cheat Sheet
Practical techniques for transforming raw data into model-ready features, including encoding, scaling, binning, and interaction terms with pandas and scikit-learn.
Encoding & Scaling
Convert categories and numeric ranges for modeling.
import pandas as pdfrom sklearn.preprocessing import OneHotEncoder, LabelEncoder, StandardScaler, MinMaxScaler# One-hot encoding (for nominal categories)df_encoded = pd.get_dummies(df, columns=["city"], drop_first=True)# Ordinal / label encoding (for ordered categories)le = LabelEncoder()df["size_encoded"] = le.fit_transform(df["size"]) # e.g. S, M, L -> 0, 1, 2# Target/mean encoding (compute within CV folds to avoid leakage!)means = df.groupby("category")["target"].mean()df["category_encoded"] = df["category"].map(means)# Scalingscaler = StandardScaler() # mean=0, std=1df[["age_scaled"]] = scaler.fit_transform(df[["age"]])minmax = MinMaxScaler() # scales to [0, 1]df[["income_scaled"]] = minmax.fit_transform(df[["income"]])
Creating New Features
Date parts, bins, and interaction terms.
import pandas as pdimport numpy as np# Date/time featuresdf["date"] = pd.to_datetime(df["date"])df["day_of_week"] = df["date"].dt.dayofweekdf["month"] = df["date"].dt.monthdf["is_weekend"] = df["day_of_week"].isin([5, 6]).astype(int)# Binning a continuous variabledf["age_bucket"] = pd.cut(df["age"], bins=[0, 18, 35, 60, 100], labels=["teen", "young_adult", "adult", "senior"])# Interaction featuresdf["price_per_sqft"] = df["price"] / df["sqft"]df["income_x_education"] = df["income"] * df["education_years"]# Log transform for right-skewed datadf["log_income"] = np.log1p(df["income"]) # log1p handles zeros safely
Encoding Techniques
Ways to turn categorical variables into numbers.
- One-hot encoding- creates a binary column per category; best for low-cardinality nominal features
- Label/ordinal encoding- maps categories to integers; only valid when order is meaningful
- Target encoding- replaces category with mean of target; must be computed within CV folds to avoid leakage
- Frequency encoding- replaces category with its occurrence count/frequency
- Hashing trick- hashes high-cardinality categories into a fixed number of buckets
- Embeddings- learned dense vectors for categories, common in deep learning pipelines
Scaling Methods
How to rescale numeric features.
- StandardScaler- centers to mean 0, std 1; assumes roughly Gaussian data, sensitive to outliers
- MinMaxScaler- rescales to a fixed range (e.g. [0,1]); preserves shape but sensitive to outliers
- RobustScaler- uses median and IQR; robust to outliers
- Normalizer- scales each sample (row) to unit norm, not each feature
- Log/Box-Cox transform- reduces right-skew and stabilizes variance
- PolynomialFeatures- generates interaction and power terms from existing features
Lag & Rolling Window Features
Engineer time-aware features for sequential/panel data without leaking the future.
import pandas as pddf = df.sort_values(["series_id", "date"])g = df.groupby("series_id")["value"]# Lag features -- shift(k) never looks aheaddf["lag_1"] = g.shift(1)df["lag_7"] = g.shift(7)# Rolling stats computed on PAST values only (shift before rolling)df["roll_mean_7"] = g.shift(1).rolling(window=7, min_periods=3).mean()df["roll_std_7"] = g.shift(1).rolling(window=7, min_periods=3).std()# Expanding stats -- cumulative up to but excluding current rowdf["expanding_mean"] = g.shift(1).expanding(min_periods=1).mean()# Diff / pct change vs prior perioddf["diff_1"] = g.diff(1)df["pct_change_1"] = g.pct_change(1)
Cyclical Encoding of Periodic Features
Encode hour/day/month as sin-cos pairs so the model sees Dec and Jan as adjacent.
import numpy as npdef add_cyclical(df, col, period): df[f"{col}_sin"] = np.sin(2 * np.pi * df[col] / period) df[f"{col}_cos"] = np.cos(2 * np.pi * df[col] / period) return dfdf["hour"] = df["timestamp"].dt.hourdf["day_of_week"] = df["timestamp"].dt.dayofweekdf["month"] = df["timestamp"].dt.monthdf = add_cyclical(df, "hour", 24)df = add_cyclical(df, "day_of_week", 7)df = add_cyclical(df, "month", 12)# Drop the raw integer column -- sin/cos pair replaces it, not augments itdf = df.drop(columns=["hour", "day_of_week", "month"])
Smoothed / K-Fold Target Encoding
Regularize target encoding with Bayesian smoothing and out-of-fold computation to curb overfitting on rare categories.
import numpy as npfrom sklearn.model_selection import KFolddef smoothed_target_encode(train, col, target, m=10): global_mean = train[target].mean() agg = train.groupby(col)[target].agg(["mean", "count"]) smooth = (agg["count"] * agg["mean"] + m * global_mean) / (agg["count"] + m) return smooth# Out-of-fold encoding avoids leaking each row's own target into its encodingkf = KFold(n_splits=5, shuffle=True, random_state=42)train["cat_te"] = np.nanfor tr_idx, val_idx in kf.split(train): mapping = smoothed_target_encode(train.iloc[tr_idx], "category", "target", m=10) train.loc[train.index[val_idx], "cat_te"] = ( train.iloc[val_idx]["category"].map(mapping).fillna(train["target"].mean()) )
Automated Feature Synthesis with Featuretools
Generate deep aggregate features across related tables via deep feature synthesis.
import featuretools as ftes = ft.EntitySet(id="orders_data")es = es.add_dataframe(dataframe_name="customers", dataframe=customers_df, index="customer_id")es = es.add_dataframe(dataframe_name="orders", dataframe=orders_df, index="order_id", time_index="order_date")es = es.add_relationship("customers", "customer_id", "orders", "customer_id")feature_matrix, feature_defs = ft.dfs( entityset=es, target_dataframe_name="customers", agg_primitives=["mean", "sum", "count", "std", "trend"], trans_primitives=["day", "month", "weekday"], max_depth=2,)
Feature Selection Methods
Techniques for cutting a feature set down to the signal that actually matters.
- Variance threshold- drops near-constant features with variance below a cutoff, cheapest first filter
- Mutual information- scores nonlinear dependence between each feature and the target (mutual_info_classif/regression)
- RFECV- recursive feature elimination with cross-validation, ranks and prunes features by model performance
- Permutation importance- shuffles one feature at a time and measures the drop in held-out score; model-agnostic
- SHAP values- game-theoretic per-feature attribution, reveals direction and magnitude of each feature's effect
- Boruta- compares real features against shadow (shuffled) copies to decide statistical relevance
- L1-based selection- Lasso/L1-penalized models drive irrelevant feature coefficients to exactly zero
- Correlation/VIF filtering- removes redundant collinear features that inflate variance without adding signal
Fit scalers and encoders only on the training fold, then transform validation/test data with those fitted parameters — fitting on the full dataset before splitting silently leaks information and inflates your validation score.