LLM Fine-Tuning Basics Cheat Sheet
Covers full fine-tuning versus parameter-efficient methods like LoRA and QLoRA, and shows how to configure PEFT for fine-tuning with Hugging Face.
Core Concepts
Approaches to adapting a pretrained LLM.
- Full fine-tuning- Updates all model weights; most expressive but requires large GPU memory and risks catastrophic forgetting
- LoRA (Low-Rank Adaptation)- Freezes the base model and trains small low-rank matrices injected into attention/linear layers, drastically cutting trainable parameters
- QLoRA- LoRA applied on top of a 4-bit quantized frozen base model, enabling fine-tuning of large models on a single GPU
- PEFT (Parameter-Efficient Fine-Tuning)- Umbrella term for methods (LoRA, prefix tuning, adapters) that train a small fraction of parameters
- Instruction tuning- Fine-tuning on (instruction, response) pairs so the model follows natural-language instructions better
- Catastrophic forgetting- Fine-tuning on a narrow task can degrade the model's general capabilities from pretraining
LoRA Fine-Tuning with PEFT
Attach LoRA adapters to a Hugging Face model.
from peft import LoraConfig, get_peft_model, TaskTypefrom transformers import AutoModelForCausalLMmodel = AutoModelForCausalLM.from_pretrained("meta-llama/Llama-3-8b")lora_config = LoraConfig( task_type=TaskType.CAUSAL_LM, r=8, # rank of the low-rank matrices lora_alpha=16, # scaling factor lora_dropout=0.05, target_modules=["q_proj", "v_proj"],)model = get_peft_model(model, lora_config)model.print_trainable_parameters() # typically < 1% of total params
Loading a 4-bit Quantized Base Model (QLoRA)
Load the frozen base model in 4-bit precision before attaching LoRA adapters.
from transformers import AutoModelForCausalLM, BitsAndBytesConfigimport torchbnb_config = BitsAndBytesConfig( load_in_4bit=True, bnb_4bit_quant_type="nf4", bnb_4bit_compute_dtype=torch.bfloat16, bnb_4bit_use_double_quant=True,)model = AutoModelForCausalLM.from_pretrained( "meta-llama/Llama-3-8b", quantization_config=bnb_config, device_map="auto")# Attach LoraConfig from above with get_peft_model(model, lora_config)
Choosing an Approach
Match the method to your compute budget and goal.
- Small dataset, limited GPU- LoRA or QLoRA; a few hours on a single consumer/cloud GPU is often enough
- Need every ounce of capability- Full fine-tuning, if you have multi-GPU compute and a large, high-quality dataset
- Just steering behavior/format- Prompt engineering or few-shot prompting first -- often cheaper than any fine-tuning
- Deploying many task variants- LoRA adapters are small (MBs) and swappable on top of one shared base model
Advanced PEFT Variants
Beyond plain LoRA -- methods that trade off parameters, quality, and flexibility differently.
- DoRA (Weight-Decomposed LoRA)- Decomposes weights into magnitude and direction, applying LoRA only to the direction component for closer-to-full-fine-tuning quality at LoRA's cost
- AdaLoRA- Allocates rank budget adaptively across layers during training instead of a fixed rank everywhere, pruning less important singular values
- IA3- Learns per-channel rescaling vectors for activations instead of low-rank matrices, training even fewer parameters than LoRA
- Prefix tuning- Prepends trainable 'virtual token' vectors to the keys/values of each attention layer instead of modifying weights
- Prompt tuning- Trains only a small set of soft prompt embeddings prepended to the input, leaving the entire model frozen
- Adapter fusion- Combines multiple task-specific adapters, e.g. via weighted averaging or a learned gate, so one base model serves several fine-tuned behaviors
Instruction Tuning with TRL's SFTTrainer
Configure a supervised fine-tuning run with sequence packing.
from trl import SFTTrainer, SFTConfigfrom datasets import load_datasetdataset = load_dataset("json", data_files="instructions.jsonl", split="train")config = SFTConfig( output_dir="./sft-out", per_device_train_batch_size=4, gradient_accumulation_steps=4, # effective batch size = 16 learning_rate=2e-4, num_train_epochs=3, max_seq_length=2048, packing=True, # pack multiple short examples per sequence)trainer = SFTTrainer( model=model, # already wrapped with get_peft_model(...) train_dataset=dataset, args=config, dataset_text_field="text",)trainer.train()
Merging LoRA Adapters for Deployment
Fold trained adapter weights into the base model so inference needs no PEFT wrapper.
from peft import PeftModelfrom transformers import AutoModelForCausalLMbase_model = AutoModelForCausalLM.from_pretrained("meta-llama/Llama-3-8b")peft_model = PeftModel.from_pretrained(base_model, "./sft-out/checkpoint-final")merged_model = peft_model.merge_and_unload() # folds LoRA deltas into base weightsmerged_model.save_pretrained("./merged-model")# Deploy merged_model directly -- no PEFT wrapper or adapter loading needed at inference
Preference Alignment with DPO
A lightweight alignment step after instruction tuning, using preference pairs instead of a reward model.
from trl import DPOTrainer, DPOConfigfrom datasets import load_dataset# Each example: prompt, chosen (preferred) response, rejected responsedataset = load_dataset("json", data_files="preferences.jsonl", split="train")config = DPOConfig( output_dir="./dpo-out", beta=0.1, # KL penalty strength vs. the reference (SFT) model learning_rate=5e-6, per_device_train_batch_size=2,)dpo_trainer = DPOTrainer( model=sft_model, # the instruction-tuned model to align ref_model=None, # None reuses a frozen copy of `model` as reference args=config, train_dataset=dataset,)dpo_trainer.train()
Hyperparameter Choices That Matter
What actually moves quality and stability for PEFT runs.
- LoRA rank (r)- Higher rank (16-64) adds capacity for complex tasks; 4-8 is often enough for narrow style/format adaptation and trains faster with less overfitting risk
- lora_alpha- Scales the adapter's contribution (effective scale = alpha/r); a common heuristic is setting alpha to roughly 2x the rank
- Learning rate- PEFT methods tolerate much higher learning rates than full fine-tuning: 1e-4 to 3e-4 for LoRA vs. 1e-5 to 5e-5 for full fine-tuning
- Effective batch size- Use gradient_accumulation_steps to reach an effective batch size of 16-64 even on a single small GPU
- target_modules- Extending beyond q_proj/v_proj to all linear layers (gate_proj, up_proj, down_proj, k_proj, o_proj) improves quality at the cost of more trainable parameters
- Epochs- 1-3 epochs is typical for instruction tuning on a few thousand examples; more risks overfitting and catastrophic forgetting on small datasets
Always keep a held-out eval set of general-purpose prompts (not just your fine-tuning task) and check it before/after fine-tuning -- a model that aces your narrow dataset but has quietly forgotten general instruction-following isn't actually an improvement.