Diffusion Models Cheat Sheet
Covers the forward noising process, learned reverse denoising process, and practical text-to-image inference with the Hugging Face diffusers library.
Core Concepts
How diffusion models generate data.
- Forward (noising) process- Fixed process that gradually adds Gaussian noise to data over T timesteps until it becomes pure noise
- Reverse (denoising) process- A neural network learns to predict and remove noise at each timestep, step by step turning noise back into data
- Noise schedule (beta_t)- Controls how much noise is added at each timestep; linear and cosine schedules are common
- Denoising objective- The model (typically a U-Net) is trained to predict the noise epsilon added at a given timestep, minimizing MSE
- Classifier-free guidance- Blends conditional and unconditional noise predictions at inference to strengthen adherence to a text prompt
- Latent diffusion- Runs the diffusion process in a compressed VAE latent space instead of pixel space for efficiency (e.g., Stable Diffusion)
Text-to-Image Inference with Diffusers
Generate an image from a text prompt using a pretrained pipeline.
from diffusers import StableDiffusionPipelineimport torchpipe = StableDiffusionPipeline.from_pretrained( "runwayml/stable-diffusion-v1-5", torch_dtype=torch.float16).to("cuda")image = pipe( prompt="a watercolor painting of a mountain lake at sunrise", num_inference_steps=30, guidance_scale=7.5,).images[0]image.save("output.png")
DDPM Reverse Sampling Step
Simplified pseudocode for one denoising step during sampling.
# x_t: noisy sample at timestep t; model predicts noise epsilonpredicted_noise = model(x_t, t)alpha_t = alphas[t]alpha_bar_t = alpha_bars[t]# Estimate the mean of p(x_{t-1} | x_t)mean = (1 / alpha_t.sqrt()) * (x_t - (1 - alpha_t) / (1 - alpha_bar_t).sqrt() * predicted_noise)if t > 0: noise = torch.randn_like(x_t) x_t_minus_1 = mean + betas[t].sqrt() * noiseelse: x_t_minus_1 = mean
Samplers / Schedulers
Different numerical methods to run the reverse process, trading speed for quality.
- DDPM- Original formulation; typically needs hundreds to ~1000 steps for high quality
- DDIM- Deterministic, non-Markovian sampler that produces good results in far fewer steps (e.g., 20-50)
- Euler / Euler Ancestral- ODE-solver-based schedulers popular for fast, high-quality sampling in Stable Diffusion pipelines
- DPM-Solver++- High-order solver that reaches good quality in as few as 15-25 steps
Advanced Concepts
Ideas beyond the basic forward/reverse formulation.
- v-prediction- Model predicts a velocity term (combination of noise and data) instead of raw noise; more numerically stable, especially with zero-terminal-SNR schedules
- Score-based SDE- Reframes diffusion as solving a stochastic differential equation over continuous time; DDPM and DDIM are discrete solvers of this same underlying process
- EMA weights- An exponential moving average of the model's weights is maintained during training and used for inference/sampling, smoothing out noisy gradient updates
- ControlNet- A trainable copy of the U-Net's encoder that injects spatial conditioning (edge maps, pose, depth) into a frozen base model without retraining it
- Latent Consistency Models (LCM)- Distilled diffusion models that learn the consistency function along the ODE trajectory, mapping noise to data in as few as 1-4 steps
- Textual inversion- Learns a new token embedding representing a concept from a handful of reference images, without touching any model weights
- Negative prompting- Encodes an undesired-concept prompt as the 'unconditional' branch of classifier-free guidance so sampling steers away from it
DDPM Training Loop
Core noise-prediction training objective with an EMA update.
from diffusers import UNet2DModel, DDPMSchedulerimport torch.nn.functional as Fscheduler = DDPMScheduler(num_train_timesteps=1000)model = UNet2DModel(sample_size=64, in_channels=3, out_channels=3)for images in dataloader: noise = torch.randn_like(images) timesteps = torch.randint( 0, scheduler.config.num_train_timesteps, (images.shape[0],), device=images.device ).long() noisy_images = scheduler.add_noise(images, noise, timesteps) noise_pred = model(noisy_images, timesteps).sample loss = F.mse_loss(noise_pred, noise) loss.backward() optimizer.step() optimizer.zero_grad() ema_model.step(model.parameters()) # update EMA weights used for sampling
Classifier-Free Guidance, Manually
How CFG actually combines conditional and unconditional predictions.
# Run the U-Net twice per step: once with the text embedding, once with an empty/null embeddingnoise_pred_uncond = unet(latents, t, encoder_hidden_states=uncond_embeddings).samplenoise_pred_text = unet(latents, t, encoder_hidden_states=text_embeddings).sample# guidance_scale > 1 extrapolates away from the unconditional prediction toward the text-conditioned oneguidance_scale = 7.5noise_pred = noise_pred_uncond + guidance_scale * (noise_pred_text - noise_pred_uncond)# For negative prompting, swap uncond_embeddings for the encoding of the negative promptlatents = scheduler.step(noise_pred, t, latents).prev_sample
Inpainting with a Mask
Regenerate only the masked region of an existing image.
from diffusers import StableDiffusionInpaintPipelineimport torchfrom PIL import Imagepipe = StableDiffusionInpaintPipeline.from_pretrained( "runwayml/stable-diffusion-inpainting", torch_dtype=torch.float16).to("cuda")init_image = Image.open("photo.png").convert("RGB")mask_image = Image.open("mask.png").convert("RGB") # white = regenerate, black = keepresult = pipe( prompt="a red brick wall", image=init_image, mask_image=mask_image, strength=0.8, # how far to deviate from init_image num_inference_steps=30,).images[0]
ControlNet Conditioned Generation
Steer output with a spatial condition (edge map) instead of text alone.
from diffusers import StableDiffusionControlNetPipeline, ControlNetModelimport torchcontrolnet = ControlNetModel.from_pretrained( "lllyasviel/sd-controlnet-canny", torch_dtype=torch.float16)pipe = StableDiffusionControlNetPipeline.from_pretrained( "runwayml/stable-diffusion-v1-5", controlnet=controlnet, torch_dtype=torch.float16).to("cuda")# canny_edge_image is a preprocessed edge map, e.g. via cv2.Canny(img, 100, 200)image = pipe( prompt="a futuristic city street", image=canny_edge_image, controlnet_conditioning_scale=1.0,).images[0]
Guidance scale is a quality/diversity trade-off, not a free lunch -- pushing it too high (beyond roughly 12-15) over-saturates images and reduces diversity even though it looks like it should 'improve' prompt adherence.