Hardware Guide

How Much VRAM Do You Need to Fine-Tune an LLM?

Last updated: October 2, 2026 — model data is refreshed automatically from the id8 dataset

Fine-tuning needs more memory than running a model, and how much more depends on the method. This page gives figures for common model sizes under QLoRA, LoRA and full fine-tuning, and explains the settings that move the number most.

Any model, method, sequence length and batch size. Free, runs in your browser.

Calculate for your GPU and model →

The numbers

Settings for this table: Unsloth, sequence length 4,096 tokens, batch size 1, gradient checkpointing on, LoRA rank 16.

ModelQLoRA (4-bit)LoRA (8-bit)LoRA (16-bit)Full fine-tune
Llama 3.2 1B2.5 GB2.8 GB3.5 GB16.4 GB
Qwen 2.5 7B7.8 GB10.7 GB16.8 GB115.6 GB
Llama 3.1 8B7.7 GB10.8 GB17.3 GB121.1 GB
Qwen 3 14B12.1 GB18.1 GB30.4 GB222.2 GB
Qwen 2.5 32B21.8 GB35.5 GB63.9 GB479.4 GB
Llama 3.3 70B42.0 GB72.6 GB135.8 GB1045.8 GB

The figures come from the same engine as Can I Fine-Tune LLM?. It reads each model's real architecture and is fitted to Unsloth's published memory benchmarks.

What the three methods mean

  • QLoRA: the base model is loaded in 4-bit and frozen. Only small adapter matrices are trained. Lowest memory.
  • LoRA (8-bit or 16-bit): the same adapters, with the base model at higher precision. More memory, slightly better fidelity.
  • Full fine-tuning: every weight is trained. Needs memory for the weights, their gradients and the optimizer's state, several times the size of the model.

More detail is in QLoRA vs LoRA vs Full Fine-Tuning.

Where the memory goes

  1. The base weights. A quarter of full size in 4-bit, half in 8-bit.
  2. Activations. Stored for the backward pass. This grows with sequence length and batch size.
  3. Adapter weights, gradients and optimizer state. Small for LoRA and QLoRA; very large for full fine-tuning.
  4. Framework overhead. The CUDA context and working buffers.

For QLoRA on a small model, activations can be the largest item, which is why the settings below matter.

Settings that change the number

SettingEffect on VRAM
Sequence lengthRoughly proportional. Halving it is the biggest single saving.
Batch sizeRoughly proportional for activations. Use batch 1 with gradient accumulation.
Gradient checkpointingLarge saving, at the cost of slower training.
LoRA rankSmall effect. Rank 8 to 32 changes little.
MethodQLoRA lowest, full fine-tuning highest.
FrameworkOptimised libraries use less than a plain setup.

Which GPU for which job

GPU memoryRealistic with QLoRA at 4,096-token sequences
8 GBModels up to about 8B, with little room to spare; shorten sequences
12 to 16 GBModels up to about 14B
24 GBThe 30B class
48 GBThe 70B class
80 GBThe 70B class with longer sequences, or 8-bit LoRA

Compare these bands with the QLoRA column in the table above for the model you have in mind, or enter your exact setup in the calculator.

If you are short of memory

  1. Reduce the sequence length to what your data needs. Measure your examples; most are shorter than you expect.
  2. Use batch size 1 and accumulate gradients over several steps.
  3. Turn on gradient checkpointing.
  4. Switch from LoRA to QLoRA.
  5. Choose a smaller base model. A well-tuned small model often beats a poorly tuned large one.
  6. Rent a larger GPU for a few hours. A fine-tuning run is short; buying hardware for it rarely pays.

Apple Silicon

MLX can fine-tune on a Mac. About 75% of unified memory is available to the GPU, so a 36 GB Mac has roughly 27 GB to work with. A worked example is in Fine-Tuning Gemma 4 on a MacBook M3 Pro.

Estimates, not guarantees

Real usage varies with library versions, the dataset and the optimizer. Leave 10 to 15% headroom, and run a few steps before a long job to confirm the peak.

Frequently asked questions

How much VRAM do I need to fine-tune an 8B model?

With QLoRA at 4,096-token sequences, batch size 1 and gradient checkpointing, Llama 3.1 8B needs about 7.7 GB. Full fine-tuning needs about 121.1 GB.

Can I fine-tune on a 24 GB GPU?

Yes, with QLoRA. A 24 GB card handles QLoRA on small and mid-size models. The table on this page shows which sizes fit.

What is the cheapest way to fine-tune?

QLoRA. It loads the base model in 4-bit and trains only small adapter matrices, so it needs the least memory of the three methods.

Why does sequence length matter so much?

Training stores activations for every token in every layer so that gradients can be computed. Doubling the sequence length roughly doubles that part of the memory.

Can I fine-tune on a Mac?

Yes, with MLX. About three quarters of a Mac's unified memory is available to the GPU.

Related