Transfer Learning Cheat Sheet
Covers feature extraction versus fine-tuning, freezing layers, and practical PyTorch code for adapting a pretrained model to a new task.
Core Concepts
The vocabulary of adapting pretrained models.
- Feature extraction- Freeze the pretrained backbone entirely and train only a new head on top of its fixed features
- Fine-tuning- Unfreeze some or all pretrained layers and continue training them, usually at a lower learning rate
- Frozen layer- A layer whose parameters are excluded from gradient updates (requires_grad = False)
- Domain shift- When the target task's data distribution differs meaningfully from the pretraining data, requiring more unfreezing
- Discriminative learning rates- Using smaller learning rates for earlier (more general) layers and larger rates for later (more task-specific) layers
Feature Extraction with a Frozen Backbone
Freeze a pretrained CNN and train only a new classification head.
import torch.nn as nnfrom torchvision import modelsmodel = models.resnet50(weights='IMAGENET1K_V2')for param in model.parameters(): param.requires_grad = False # freeze everything# Replace the final layer - new params are trainable by defaultmodel.fc = nn.Linear(model.fc.in_features, num_classes)optimizer = torch.optim.Adam(model.fc.parameters(), lr=1e-3)
Fine-Tuning the Last Few Layers
Gradually unfreeze layers closest to the output for a more task-specific adaptation.
# Unfreeze just layer4 and the classifier headfor name, param in model.named_parameters(): param.requires_grad = name.startswith('layer4') or name.startswith('fc')optimizer = torch.optim.Adam([ {'params': model.layer4.parameters(), 'lr': 1e-5}, # smaller lr for pretrained layers {'params': model.fc.parameters(), 'lr': 1e-3}, # larger lr for new head])
Choosing a Strategy
How much of the model to adapt.
- Small dataset, similar domain- Feature extraction (freeze everything) usually works best and avoids overfitting
- Large dataset, similar domain- Fine-tune the whole network at a low learning rate
- Small dataset, different domain- Fine-tune only the last few layers; early layers capture generic features (edges, textures) that transfer well
- Large dataset, different domain- Fine-tune the whole network, or train from scratch if the domain gap is extreme
Gradual Unfreezing Across Epochs
Unfreeze one block at a time so earlier layers don't get destabilized by large early gradients.
layer_groups = [model.layer1, model.layer2, model.layer3, model.layer4]# Start fully frozen except the headfor group in layer_groups: for p in group.parameters(): p.requires_grad = Falsefor epoch in range(num_epochs): # unfreeze one more group every few epochs, working backward from the output if epoch in (3, 6, 9, 12): idx = {3: -1, 6: -2, 9: -3, 12: -4}[epoch] for p in layer_groups[idx].parameters(): p.requires_grad = True optimizer = torch.optim.Adam( filter(lambda p: p.requires_grad, model.parameters()), lr=1e-4 ) train_one_epoch(model, optimizer)
Auditing Which Parameters Will Actually Update
A quick sanity check to catch layers that were accidentally left frozen (or unfrozen).
def report_trainable(model): total, trainable = 0, 0 for name, p in model.named_parameters(): total += p.numel() if p.requires_grad: trainable += p.numel() print(f"trainable: {name:40s} {tuple(p.shape)}") pct = 100 * trainable / total print(f"{trainable:,} / {total:,} params trainable ({pct:.2f}%)")report_trainable(model)
Beyond Freeze/Unfreeze
Parameter-efficient and regularized alternatives to full fine-tuning.
- Layer-wise LR decay (LLRD)- Apply a multiplicative decay factor (e.g. 0.9) to the learning rate per layer going backward from the output, common when fine-tuning transformers
- BitFit- Fine-tune only the bias terms of the pretrained network, leaving weight matrices frozen -- surprisingly competitive with far fewer trainable params
- Adapters / LoRA- Insert small trainable bottleneck modules (or low-rank updates) alongside frozen pretrained weights instead of updating the originals directly
- Elastic Weight Consolidation (EWC)- Adds a penalty proportional to the Fisher information that discourages moving weights away from their pretrained values, mitigating catastrophic forgetting
- Warmup for the new head- Train only the newly initialized head for a few epochs before unfreezing the backbone, so early noisy gradients don't propagate into pretrained weights
- Discriminative fine-tuning + slanted triangular LR- Combine per-layer learning rates with a schedule that rises quickly then decays slowly (used in ULMFiT-style NLP transfer learning)
Layer-wise Learning Rate Decay for a Transformer Encoder
Assign progressively smaller learning rates to earlier encoder layers.
def llrd_param_groups(model, base_lr=2e-5, decay=0.9, head_lr=1e-3): groups = [{'params': model.classifier.parameters(), 'lr': head_lr}] layers = list(model.encoder.layer)[::-1] # reverse: last layer first lr = base_lr for layer in layers: groups.append({'params': layer.parameters(), 'lr': lr}) lr *= decay groups.append({'params': model.embeddings.parameters(), 'lr': lr}) return groupsoptimizer = torch.optim.AdamW(llrd_param_groups(model))
Precomputing Frozen Features for a Downstream Classifier
When the backbone is fully frozen, extract features once and train a lightweight classifier offline -- much faster than re-running the backbone every epoch.
import torchbackbone.eval()features, labels = [], []with torch.no_grad(): for x, y in dataloader: feats = backbone(x.cuda()).flatten(1) # e.g. (B, 2048) for ResNet50 features.append(feats.cpu()) labels.append(y)X = torch.cat(features)y = torch.cat(labels)# Now fit any classifier (sklearn LogisticRegression, SVM, or a small MLP) on (X, y)
Use a much smaller learning rate for unfrozen pretrained layers than for a newly initialized head -- a single high learning rate applied to both can quickly destroy useful pretrained weights ('catastrophic forgetting').