Guide

QLoRA vs LoRA vs Full Fine-Tuning: Differences, Memory and When to Use Each

Last updated: October 2, 2026 — model data is refreshed automatically from the id8 dataset

There are three common ways to fine-tune a language model, and they differ by an order of magnitude in the hardware they need. For most projects the cheapest one is also the right one. This guide explains why, and when it is not.

QLoRA, LoRA and full fine-tuning for any model. Free.

Check what your GPU can fine-tune →

The three methods in one paragraph each

Full fine-tuning updates every weight in the model. It has the most capacity to change behaviour and needs the most memory: the weights, a gradient for each weight, and the optimizer's running statistics.

LoRA, low-rank adaptation, freezes the model and adds a pair of small matrices beside chosen layers. Only those matrices are trained. They hold a tiny fraction of the model's parameters, so gradients and optimizer state are small. The result is saved as an adapter file of a few megabytes to a few hundred.

QLoRA is LoRA with the frozen base model loaded in 4-bit precision. The adapters are still trained at normal precision. Because the base weights are the largest item in memory, this cuts the requirement dramatically.

Memory for the same model

Sequence length 4,096, batch size 1, gradient checkpointing on, rank 16:

ModelQLoRA (4-bit)LoRA (8-bit)LoRA (16-bit)Full fine-tune
Llama 3.1 8B7.7 GB10.8 GB17.3 GB121.1 GB
Qwen 3 14B12.1 GB18.1 GB30.4 GB222.2 GB
Qwen 2.5 32B21.8 GB35.5 GB63.9 GB479.4 GB

Read across a row to see what the method alone costs you. The same model that trains on a consumer card with QLoRA can need data-centre hardware for a full fine-tune.

Side by side

QLoRALoRAFull fine-tuning
Base model precision4-bit8-bit or 16-bit16-bit or 32-bit
Weights trainedAdapters onlyAdapters onlyAll
VRAMLowestMediumHighest
Training speed per stepSlower than LoRAFastSlowest
OutputSmall adapter fileSmall adapter fileA full copy of the model
Risk of forgetting what the model knewLowLowHigher
Capacity to learn something very newModerateModerateHighest

How to choose

Start with QLoRA when:

  • your GPU has 24 GB or less;
  • the task is narrow: a tone of voice, a document format, a classification scheme, domain vocabulary;
  • you have hundreds to tens of thousands of examples.

Use LoRA at higher precision when you have memory to spare and want to remove quantization as a variable, or when training speed matters more than memory.

Consider a full fine-tune when:

  • you have a large dataset, hundreds of thousands of examples or more;
  • you are teaching something the base model lacks broadly, such as a new language or a new modality;
  • you have the hardware and time, and can test for regressions on what the model used to do well.

What matters more than the method

  • Data quality. A few thousand clean, consistent examples beat a large noisy set. Most disappointing fine-tunes are data problems.
  • A held-out test set. Keep 5 to 10% of examples out of training and measure on them.
  • The baseline. Try a good prompt with examples first. If prompting gets you there, you do not need to fine-tune.
  • Learning rate and epochs. Too many passes over a small dataset makes the model repeat its training examples.

After training

LoRA and QLoRA produce an adapter. You can load it on top of the base model at run time, or merge it into the weights and quantize the result for local use. One base model can serve several adapters for different tasks, which is far cheaper to store than several full models.

Hardware check

Before you start, enter your model, method and sequence length in Can I Fine-Tune LLM?. For the settings that move the number most, see How Much VRAM Do You Need to Fine-Tune an LLM?.

Frequently asked questions

What is the difference between LoRA and QLoRA?

Both train small adapter matrices and leave the base model frozen. QLoRA also loads the base model in 4-bit precision, which cuts memory sharply.

Is QLoRA worse than LoRA?

The original QLoRA paper reported results close to full-precision fine-tuning on its tests. In practice the difference is small for most tasks, and the memory saving is large.

When should I do a full fine-tune?

When you have a large dataset, need to change the model's behaviour deeply, for example teaching a new language, and have the hardware. For adapting style or a narrow task, adapters are enough.

How much memory does each method need for an 8B model?

For Llama 3.1 8B at 4,096-token sequences: QLoRA about 7.7 GB, 16-bit LoRA about 17.3 GB, full fine-tuning about 121.1 GB.

What LoRA rank should I use?

8 to 32 suits most tasks. Start at 16. A higher rank adds capacity and a little memory; it does not fix a poor dataset.

Related