LoRA vs QLoRA: Which Fine-Tuning Method Should I Use?
LoRA trains small adapter matrices while keeping the pretrained weights frozen. QLoRA applies that approach to a base model stored in 4-bit precision. Start with LoRA when the base model fits comfortably; consider QLoRA when its frozen weights consume too much training memory.
Both methods still need memory for activations, adapter gradients, and optimizer state. Larger batches and longer sequences increase activation and temporary-buffer memory. For a fixed adapter setup, gradient and optimizer-state sizes depend on the trainable parameters.
The difference is the base weights
LoRA represents a weight update with two small matrices rather than training the whole weight matrix. If the original weight has dimensions d × k, an adapter of rank r trains r × k and d × r matrices. Rank determines the adapter’s size; it does not supply a universal measure of task difficulty.
| Choice | Frozen base | Trainable part | Main trade-off |
|---|---|---|---|
| LoRA | Commonly BF16 or FP16 | Low-rank adapters | Larger base-weight allocation |
| QLoRA | 4-bit quantized weights | Low-rank adapters, with higher precision arithmetic | Less base-weight memory, with dequantization work |
The original LoRA paper describes frozen weights and trainable low-rank updates. An adapter can be merged into compatible base weights after training, avoiding a separate adapter operation during inference. Quantized deployment and merging require support from the selected tooling.
The QLoRA paper used NF4, a 4-bit representation designed for normally distributed weights, plus double quantization and paged optimizers. It fine-tuned a 65B model on one 48 GB GPU in its tested setup. That result does not promise the same fit for another model, sequence length, or training configuration.
Compare one controlled training run
Keep the base checkpoint, examples, adapter targets, rank, sequence length, and evaluation set constant. Record peak memory, training time, and held-out task quality. If QLoRA saves memory but slows the run, decide whether fitting on fewer GPUs compensates for that time.
Also test the actual inference artifact. Adapter quality measured with the training stack may differ after merging or further quantization.
The LoRA and QLoRA section of the LLM Engineering Guide explains the parameter update and its hardware implications.