Computer Vision Basics Cheat Sheet
Core computer vision concepts and workflows, covering image preprocessing, convolutional filters, and building a basic CNN classifier with PyTorch.
Image Loading & Preprocessing
Load, resize, and augment images.
import cv2import torchvision.transforms as T# Load and inspect an image with OpenCV (BGR order by default)img = cv2.imread("photo.jpg")img_rgb = cv2.cvtColor(img, cv2.COLOR_BGR2RGB)gray = cv2.cvtColor(img, cv2.COLOR_BGR2GRAY)# Resize, blur, edge detectionresized = cv2.resize(img_rgb, (224, 224))blurred = cv2.GaussianBlur(img_rgb, (5, 5), sigmaX=0)edges = cv2.Canny(gray, threshold1=100, threshold2=200)# torchvision transforms for a training pipelinetransform = T.Compose([ T.Resize((224, 224)), T.RandomHorizontalFlip(p=0.5), T.ToTensor(), T.Normalize(mean=[0.485, 0.456, 0.406], std=[0.229, 0.224, 0.225]),])
A Simple CNN
Convolution, pooling, and classification layers.
import torch.nn as nnclass SimpleCNN(nn.Module): def __init__(self, num_classes=10): super().__init__() self.conv1 = nn.Conv2d(3, 32, kernel_size=3, padding=1) self.conv2 = nn.Conv2d(32, 64, kernel_size=3, padding=1) self.pool = nn.MaxPool2d(2, 2) self.fc1 = nn.Linear(64 * 56 * 56, 128) self.fc2 = nn.Linear(128, num_classes) self.relu = nn.ReLU() def forward(self, x): x = self.pool(self.relu(self.conv1(x))) # 224 -> 112 x = self.pool(self.relu(self.conv2(x))) # 112 -> 56 x = x.view(x.size(0), -1) x = self.relu(self.fc1(x)) return self.fc2(x)model = SimpleCNN(num_classes=10)
Computer Vision Concepts
Core building blocks of CNN-based vision models.
- Convolution- sliding a learnable filter/kernel over an image to detect local patterns like edges
- Pooling- downsamples feature maps (e.g. max pooling) to reduce spatial size and add translation invariance
- Stride- step size the filter moves each time; larger stride reduces output size
- Padding- adds border pixels so output size can match input size ('same' padding)
- Feature map- the output of applying a convolutional filter to the input
- Data augmentation- random flips/rotations/crops applied during training to improve generalization
- Transfer learning- reusing a model pretrained on a large dataset (e.g. ImageNet) and fine-tuning it on your task
- IoU (Intersection over Union)- overlap metric between predicted and ground-truth bounding boxes, used in object detection
Common CV Tasks
Typical problems solved with computer vision.
- Image classification- assigning a single label to an entire image
- Object detection- locating and classifying multiple objects with bounding boxes (e.g. YOLO, Faster R-CNN)
- Semantic segmentation- classifying every pixel in an image into a category
- Instance segmentation- segmenting individual object instances, distinguishing overlapping objects of the same class
- Image generation- synthesizing new images (e.g. GANs, diffusion models)
Transfer Learning with a Pretrained Backbone
Fine-tune a pretrained ResNet by freezing early layers and replacing the classifier head.
import torchimport torch.nn as nnfrom torchvision.models import resnet50, ResNet50_Weightsmodel = resnet50(weights=ResNet50_Weights.IMAGENET1K_V2)# Freeze the convolutional backbone, only train the new headfor param in model.parameters(): param.requires_grad = Falsenum_features = model.fc.in_featuresmodel.fc = nn.Linear(num_features, num_classes := 5)optimizer = torch.optim.AdamW(model.fc.parameters(), lr=1e-3)# After a few epochs, unfreeze the last block for fine-tuning at a lower LRfor param in model.layer4.parameters(): param.requires_grad = Trueoptimizer.add_param_group({"params": model.layer4.parameters(), "lr": 1e-5})
Object Detection with torchvision
Load a pretrained Faster R-CNN and run inference to get boxes, labels, and scores.
import torchfrom torchvision.models.detection import fasterrcnn_resnet50_fpn_v2, FasterRCNN_ResNet50_FPN_V2_Weightsfrom torchvision.transforms.functional import to_tensorweights = FasterRCNN_ResNet50_FPN_V2_Weights.DEFAULTmodel = fasterrcnn_resnet50_fpn_v2(weights=weights)model.eval()img_tensor = to_tensor(img_rgb) # [C, H, W], values in [0, 1]with torch.no_grad(): predictions = model([img_tensor])[0]boxes = predictions["boxes"] # [N, 4] xyxy formatlabels = predictions["labels"] # class indicesscores = predictions["scores"] # confidence per box# Keep only confident detectionskeep = scores > 0.5boxes, labels, scores = boxes[keep], labels[keep], scores[keep]
Non-Max Suppression & mAP Evaluation
Deduplicate overlapping boxes and score detections with mean average precision.
import torchfrom torchvision.ops import nms, box_iou# Non-max suppression: drop lower-confidence boxes that overlap a kept box past iou_thresholdkeep_idx = nms(boxes, scores, iou_threshold=0.5)boxes, scores, labels = boxes[keep_idx], scores[keep_idx], labels[keep_idx]# IoU matrix between predicted and ground-truth boxesious = box_iou(boxes, gt_boxes)# mAP with torchmetrics aggregates precision-recall across IoU thresholds and classesfrom torchmetrics.detection.mean_ap import MeanAveragePrecisionmetric = MeanAveragePrecision(iou_type="bbox")metric.update( preds=[{"boxes": boxes, "scores": scores, "labels": labels}], target=[{"boxes": gt_boxes, "labels": gt_labels}],)print(metric.compute()) # includes map, map_50, map_75
Advanced Augmentation with Albumentations
Compose fast, GPU-friendly augmentations that also transform bounding boxes/masks in sync.
import albumentations as Afrom albumentations.pytorch import ToTensorV2transform = A.Compose([ A.RandomResizedCrop(size=(224, 224), scale=(0.8, 1.0)), A.HorizontalFlip(p=0.5), A.ColorJitter(brightness=0.2, contrast=0.2, saturation=0.2, hue=0.1, p=0.5), A.CoarseDropout(num_holes_range=(1, 8), hole_height_range=(8, 16), hole_width_range=(8, 16), p=0.3), A.Normalize(mean=[0.485, 0.456, 0.406], std=[0.229, 0.224, 0.225]), ToTensorV2(),], bbox_params=A.BboxParams(format="pascal_voc", label_fields=["class_labels"]))augmented = transform(image=img_rgb, bboxes=boxes, class_labels=labels)img_t, boxes_aug, labels_aug = augmented["image"], augmented["bboxes"], augmented["class_labels"]
Advanced CV Concepts
Terms that appear once you move beyond a single classifier CNN.
- Batch normalization- normalizes layer activations per mini-batch, stabilizing and speeding up training of deep CNNs
- Receptive field- the region of the input image that influences a given neuron's activation; grows with network depth and stride
- Feature Pyramid Network (FPN)- combines multi-scale feature maps so detectors can find both small and large objects
- Anchor boxes- predefined reference boxes of varying scale/aspect ratio that detectors like Faster R-CNN and YOLO regress offsets from
- Non-max suppression (NMS)- post-processing step that removes duplicate overlapping detections of the same object
- Grad-CAM- gradient-based visualization highlighting which image regions most influenced a CNN's prediction
- Mixed precision training- using fp16/bf16 for most ops while keeping fp32 master weights, cutting memory use and speeding up training
- ONNX export- converting a trained model to a framework-agnostic graph format for optimized cross-platform inference
Always normalize input images using the same mean/std the pretrained backbone was trained with (e.g. ImageNet's [0.485, 0.456, 0.406] / [0.229, 0.224, 0.225]) when fine-tuning — mismatched normalization silently degrades transfer learning performance.