LLM Fine-Tuning VRAM Calculator

Can I Fine-Tune LLM is a free VRAM estimator for QLoRA, LoRA, and full fine-tuning — 95 models, RTX 5090/4090 and Apple Silicon, 100% in your browser.

Estimate GPU memory for QLoRA, LoRA & Full Fine-Tuning of open-source LLMs

Configuration

GB
2GB192GB
64 GB
4GB256GB

4,096 tokens
512128K
1
132
COMFORTABLE
TIGHT
OOM

Top Picks for Your Hardware

Training Compatibility — All Models

Quick Reference: VRAM Needed for Local Fine-Tuning

Estimated GPU memory for common model sizes. Unsloth, batch size 1, 4K-token sequences, gradient checkpointing on.

Model Size QLoRA (4-bit) LoRA (8-bit) LoRA (16-bit) Full Fine-Tune Min GPU (QLoRA)
1B ~2.5 GB ~2.8 GB ~3.5 GB ~16.4 GB RTX 4060 (8GB)
7B ~7.8 GB ~10.7 GB ~16.8 GB ~115.6 GB RTX 3060 (12GB)
8B ~7.7 GB ~10.8 GB ~17.3 GB ~121.1 GB RTX 3060 (12GB)
14B ~12.1 GB ~18.1 GB ~30.4 GB ~222.2 GB RTX 4060 Ti (16GB)
32B ~21.8 GB ~35.5 GB ~63.9 GB ~479.4 GB RTX 4090 (24GB)
70B ~42.0 GB ~72.6 GB ~135.8 GB ~1,046 GB RTX 6000 Ada (48GB)

Unsloth, LoRA rank 16 on all linear layers, batch 1, 4,096-token sequences, gradient checkpointing on. Rows use real architectures (1B = Llama 3.2 1B, 7B = Qwen 2.5 7B, 8B = Llama 3.1 8B, 14B = Qwen 3 14B, 32B = Qwen 2.5 32B, 70B = Llama 3.3 70B). Plain Hugging Face + PEFT needs more — mostly for the full-vocabulary logits — and memory grows with every extra token of context: use the calculator above for your exact setup. Estimates are fitted to Unsloth's published context-length benchmarks (8B and 70B, ±10%); small models under Hugging Face/MLX are likely overestimated.

How we calculate?

Fine-tuning VRAM = Model Weights + Adapter Optimizer States + Activations × sequence length + Logits + Framework Overhead, computed from each model's real architecture (its config.json) and fitted to Unsloth's published context-length benchmarks. QLoRA quantizes base weights to 4-bit (NF4) and only trains small LoRA adapter matrices. LoRA keeps the base frozen at 16-bit and trains adapters. Full fine-tuning updates all weights, requiring AdamW optimizer states (8 bytes/param) and full gradients (4 bytes/param). For MoE models, ALL expert weights are trained — not just the active fraction.

What is Gradient Checkpointing?

Instead of storing all intermediate activations during the forward pass, gradient checkpointing keeps each layer's input and recomputes the rest during the backward pass. This saves most of the activation memory (not the logits) at the cost of ~20–30% slower training. The output is mathematically identical — no quality loss. Always recommended for local training on consumer GPUs.

QLoRA vs LoRA vs Full FT

QLoRA (4-bit): Best for consumer GPUs. ~0.5 bytes/param base. Trains 1–5% of parameters via adapters.
LoRA (8-bit): Middle ground. ~1 byte/param base. Slightly better quality than QLoRA.
LoRA (16-bit): bf16 base, 2 bytes/param. Best adapter quality.
Full FT: Updates all weights. ~16 bytes/param. Enterprise-only.
All methods add activation memory per token of context — see the calculator.

QLoRA vs LoRA vs Full Fine-Tuning: VRAM Requirements Compared

The fine-tuning method you choose has a 5–10× impact on VRAM consumption — which directly determines whether you can run a given model on consumer hardware at all. QLoRA, LoRA, and full fine-tuning occupy three very different points on the memory–quality tradeoff curve. Understanding where each sits lets you make the right call before you spin up a training run and hit an OOM error two hours in.

Method Base Model Precision Adapter Precision VRAM for 7B Model VRAM for 13B VRAM for 70B Notes
Full Fine-Tuning FP32 N/A ~112 GB ~208 GB ~1.12 TB Needs multi-GPU for anything >3B
Full Fine-Tuning BF16 N/A ~56 GB ~104 GB ~560 GB Still impractical without enterprise GPU
LoRA (FP16 base) FP16 FP32 adapters ~16–18 GB ~28–32 GB ~140–160 GB Popular for 24 GB consumer GPUs
QLoRA (4-bit base) NF4 4-bit BF16 adapters ~6–8 GB ~12–14 GB ~48–56 GB Consumer-grade fine-tuning
QLoRA + GC NF4 4-bit + GC BF16 adapters ~5–7 GB ~10–12 GB ~40–48 GB GC = gradient checkpointing

QLoRA achieves its dramatic memory savings through NF4 (NormalFloat4) quantization applied to the frozen base model weights. NF4 is not the same as INT4 — it uses a quantization grid that is information-theoretically optimal for normally distributed weights, which is what you find in pretrained LLMs. Because the base model is frozen and only the small LoRA adapter matrices are updated, the adapter gradients are kept in full BF16 precision to preserve training stability. The quantization error introduced by NF4 is partially compensated by double quantization, where the quantization constants themselves are quantized a second time, saving an additional 0.37 bits per parameter.

LoRA without quantization is worth the extra VRAM in two situations. First, when convergence speed matters — FP16 base models typically converge 10–20% faster than QLoRA because the optimizer sees full-precision gradients throughout. Second, when final model quality is the priority: on tasks that demand very precise domain adaptation (e.g., legal or medical text), LoRA consistently scores 1–3 percentage points higher than QLoRA on downstream benchmarks. If you have a 24 GB GPU and are running a 7B model, LoRA fits comfortably and the quality gains are usually worth it. For anything larger, QLoRA is the only viable path without multi-GPU setup.

Gradient Checkpointing: Trading Training Speed for VRAM Savings

During a standard forward pass, the model stores every intermediate activation tensor in VRAM so the backward pass can use them to compute gradients. For a 7B model with a typical batch size and sequence length, this activation memory can add 5–15 GB on top of the model weights themselves — easily pushing a workload over the edge on a 16 GB or 24 GB card. Gradient checkpointing (GC) solves this by discarding most intermediate activations as they are computed and recomputing them on-the-fly during the backward pass. The result is a 30–60% reduction in peak VRAM usage from activations, which can be the difference between fitting a model on consumer hardware and not fitting it at all.

The cost of this VRAM saving is training time. Because activations must be recomputed, gradient checkpointing typically adds 20–40% to total training time compared to a run without it. In practice, the recommendation is straightforward: always enable gradient checkpointing when fine-tuning on consumer hardware (RTX 4090 and below, or Apple M-series with 24 GB or less). The time penalty is acceptable when it unlocks the ability to train at all. Disable gradient checkpointing only when VRAM is genuinely plentiful — for example, on an A100 80 GB or H100 with a model that already fits with headroom to spare and where throughput is more important than memory frugality.

Enabling gradient checkpointing in Hugging Face Trainer is a single flag. The use_reentrant=False setting is strongly recommended when using QLoRA because the reentrant autograd checkpoint implementation has known incompatibilities with the bitsandbytes quantization hooks that QLoRA relies on:

# Hugging Face Trainer — enable gradient checkpointing
training_args = TrainingArguments(
    gradient_checkpointing=True,
    gradient_checkpointing_kwargs={"use_reentrant": False},  # recommended for QLoRA
    per_device_train_batch_size=1,
    gradient_accumulation_steps=4,  # effective batch = 4
)

The gradient_accumulation_steps=4 setting above is a complementary technique: it simulates a larger effective batch size (4 in this case) without storing more activations simultaneously, because only one micro-batch is live in memory at any given moment. Combining gradient checkpointing with gradient accumulation is the standard playbook for fine-tuning large models on consumer hardware, and both are supported natively by TRL's SFTTrainer without any additional code.

Fine-Tuning on Apple Silicon: MLX vs MPS (PyTorch)

Apple M-series chips use a unified memory architecture where the GPU and CPU share the same physical RAM pool. This is an advantage for LLM fine-tuning, with one catch: by default macOS lets the GPU use only about 75% of it (Metal's working-set limit) — about 13.5 GB on an 18 GB Mac or 36 GB on a 48 GB M4 Max — and the OS and your apps need the rest. The limit can be raised with sudo sysctl iogpu.wired_limit_mb, at the risk of starving macOS. Even so, a 48 GB Mac gives a far bigger "GPU" than any consumer NVIDIA card.

Chip Unified Memory Max Model for QLoRA (short sequences) MLX Support MPS (PyTorch) Support Practical Throughput
M3 / M4 (base) 16 GB 7B Q4 ✓ ✓ (limited) ~150–200 tok/s training
M3 Pro / M4 Pro 18–24 GB 7B–13B Q4 ✓ ✓ ~200–350 tok/s
M3 Max / M4 Max 36–48 GB 13B–34B Q4 ✓ ✓ ~400–700 tok/s
M2 Ultra / M3 Ultra 64–192 GB Up to 70B Q4 ✓ ✓ ~800–1500 tok/s

MLX is the preferred framework for fine-tuning on Apple Silicon over PyTorch's MPS (Metal Performance Shaders) backend for several reasons. MLX is written from the ground up for Apple's Metal GPU API and uses native Metal ops throughout the entire compute graph. PyTorch's MPS backend, by contrast, was originally designed for CUDA and uses a translation layer that introduces overhead, particularly for non-standard ops that fall back to CPU. In practice, MLX achieves 20–40% higher token throughput than PyTorch MPS on the same Apple Silicon hardware, and it avoids the MPS memory fragmentation bugs that can cause out-of-memory errors mid-run even when total memory use is within budget.

Fine-tuning a model with MLX requires just two steps after installing the package. Install with pip install mlx-lm, then launch a LoRA fine-tuning run with a single command: mlx_lm.lora --model mistralai/Mistral-7B-v0.1 --train --data ./data/. MLX will automatically download the model from Hugging Face, apply LoRA adapters, and train using the JSONL files in your ./data/ directory (which should contain train.jsonl and optionally valid.jsonl). Adapter weights are saved as adapters.npz and can be merged back into the base model with mlx_lm.fuse for inference.

MLX does have meaningful limitations to plan around. It does not support multi-GPU or multi-node training — all compute happens on the single unified memory pool of one machine. As of the MLX 0.x series, Flash Attention is not implemented, which means memory usage for long sequences scales quadratically rather than linearly and very long context fine-tuning (32K+ tokens) is not practical. Not all model architectures are supported yet; Llama, Mistral, Phi, Qwen, and Gemma families work well, but newer or less popular architectures may require a custom conversion. For these cases, falling back to PyTorch MPS or a cloud GPU is the practical path.

Local vs Cloud Fine-Tuning: Cost Breakdown for 2026

For a single fine-tuning run on a 7B model with 10,000 training examples — roughly 3–4 hours of wall-clock training time on a single high-end consumer GPU — the economics vary dramatically depending on whether you own your hardware or rent cloud compute. The table below compares the most common options available in 2026, with prices in both INR and USD for reference.

Option Setup Cost Cost per Run Best For Caveats
RTX 4090 (owned) ₹1.5–2L one-time ₹0 (electricity only, ~₹20–40) Frequent fine-tuning High upfront cost
MacBook M4 Max 48 GB ₹3.5–4.5L one-time ₹0 (electricity, ~₹10–20) Mac-only workflows Slower than CUDA for LLMs
RunPod (RTX 4090) ₹0 ~₹120–200/run ($1.5–2.5/hr × 3–4 hr) Occasional runs Cold start, data upload
Vast.ai (RTX 3090) ₹0 ~₹60–100/run ($0.7–1.2/hr × 3–4 hr) Budget runs Less reliable uptime
Lambda Labs (A100 80 GB) ₹0 ~₹500–700/run ($2–2.5/hr × 3–4 hr) Large models >34B Enterprise grade
Google Colab Pro+ ~₹1700/month Shared with subscription Students, hobbyists Session time limits, slow disk

The break-even calculation for owning versus renting depends primarily on how many fine-tuning runs you plan to do per year. At RunPod RTX 4090 prices of ~₹150 per run, an RTX 4090 card costing ₹1.8L amortizes after roughly 50–70 runs — which is less than one run per week over a year. If you are actively iterating on models (multiple runs per week for hyperparameter search, dataset experiments, or production model updates), owning the hardware pays off within a few months. If you run fine-tuning only a handful of times per quarter, renting cloud GPUs is almost always cheaper when you factor in electricity, cooling, and the opportunity cost of capital tied up in hardware.

Data privacy is the other major factor that the raw cost numbers do not capture. Cloud fine-tuning platforms — RunPod, Vast.ai, Lambda Labs, and Colab — all require you to upload your training dataset to their infrastructure. For general-purpose or open-domain fine-tuning this is usually acceptable. But if your training data contains proprietary business information, customer data, medical records, legal documents, or anything regulated under GDPR, HIPAA, or similar frameworks, uploading it to a third-party cloud platform creates compliance and legal exposure that may outweigh the cost savings entirely. In those cases, local fine-tuning — whether on an owned GPU or Apple Silicon — is not just a preference but a requirement. The RTX 4090 or M4 Max starts to look very cost-effective the moment data sovereignty becomes a hard constraint.

Dataset Preparation: The Step That Makes or Breaks Fine-Tuning

Even with perfect hardware and the right fine-tuning method, the output quality of your fine-tuned model is bounded by the quality of your training data. This is the part of the pipeline that practitioners underestimate most often. A model fine-tuned on noisy, inconsistent, or poorly formatted data will perform worse than the base model on your target task — a phenomenon sometimes called "fine-tuning degradation." Getting dataset preparation right is therefore not optional; it is the core engineering challenge of supervised fine-tuning.

Format: Instruction-Tuning vs Completion

The two dominant data formats for fine-tuning are instruction-response pairs and completion-only sequences. Instruction-response (also called chat or instruction-tuning format) structures each example as a system prompt plus a user message plus an expected model response. This format aligns with how modern chat models are trained and is the right choice when you want to teach the model to follow instructions in a specific style or domain. Completion-only format is simpler: it is a continuous block of text that the model learns to continue, and it is appropriate for tasks like code generation or domain-specific text continuation where instruction-following behaviour is already present in the base model and you just want to shift the output distribution.

Volume and Quality Thresholds

The minimum viable dataset size for meaningful fine-tuning is roughly 500–1000 high-quality examples for a narrow, well-defined task (e.g., reformatting product descriptions into a specific JSON schema). For broader behavioural changes — teaching a model a new domain, persona, or reasoning style — you typically need 5,000–50,000 examples. The quality threshold matters far more than quantity: 500 carefully curated, consistent, human-verified examples routinely outperform 10,000 examples scraped from the web without filtering. A practical rule of thumb is to hand-review at least 5% of your dataset before training — if you find errors in more than 5% of reviewed examples, the whole dataset likely needs another cleaning pass.

Deduplication and Train/Validation Split

Duplicate examples in a fine-tuning dataset are more damaging than in large-scale pretraining because the training corpus is much smaller and the model is more susceptible to memorising repeated patterns. Always deduplicate at the level of both exact matches and near-duplicates (using MinHash or a similar approximate deduplication method). Reserve 5–10% of your data as a held-out validation set and monitor validation loss during training — if training loss keeps falling but validation loss starts rising, the model is overfitting to your training examples and you should stop early. A train/validation split ratio of 90/10 is standard; using a completely separate test set for final evaluation (not touched during training or hyperparameter selection) is best practice for production deployments.

Data Format for MLX and Hugging Face

Both MLX and Hugging Face TRL expect training data as JSONL files where each line is a JSON object. For instruction-tuning with Hugging Face TRL's SFTTrainer, the standard format is a messages field containing a list of role/content dictionaries, matching the OpenAI chat format. For MLX, each line should have a text field with the fully formatted prompt already applied. Converting between formats is straightforward, but applying the correct chat template (using the tokenizer's apply_chat_template method) is essential — using the wrong template or formatting prompts manually without the template often introduces invisible whitespace or special token differences that silently degrade model quality at inference time.

Sequence Length and Packing

Maximum sequence length is one of the most impactful hyperparameters for both VRAM usage and training throughput. Setting it too high wastes VRAM on padding tokens and reduces the number of examples that fit per batch; setting it too low truncates examples and loses information. A practical approach is to compute the 95th-percentile token length across your dataset and set max_seq_length to that value, then review the 5% of examples that are truncated to decide if they are critical. For most instruction-tuning datasets with short to medium-length responses, a value of 1024–2048 tokens strikes the right balance between VRAM efficiency and data completeness. If your dataset has many short examples, enabling sequence packing — concatenating multiple short examples into a single training sequence separated by EOS tokens — can dramatically improve GPU utilisation and reduce training time by 30–50% without affecting model quality.

LLM Fine-Tuning FAQ & VRAM Guide

What is the difference between QLoRA and LoRA VRAM usage?

QLoRA stores the frozen base in 4-bit NF4 (~0.5 bytes/param) and trains small LoRA adapters. LoRA 16-bit keeps the base in bf16 (2 bytes/param). Full fine-tuning needs ~16 bytes/param (weights, gradients, fp32 master copy and Adam states). On top of that, every method needs activation memory that grows with sequence length — often the part that decides whether a run fits.

How much VRAM to fine-tune Llama 3 8B?

At 4,096-token sequences with gradient checkpointing: QLoRA needs about 7.7 GB with Unsloth or 19.4 GB with plain Hugging Face + PEFT; LoRA 16-bit about 17.3 GB (Unsloth); full fine-tuning about 121.1 GB. Context length matters as much as model size: on a 24 GB GPU, QLoRA fits roughly 81K tokens with Unsloth but only about 5.9K with plain Hugging Face.

Does batch size affect VRAM during fine-tuning?

Yes, significantly. Activation memory scales linearly with batch size. Doubling the batch size roughly doubles activation memory. For consumer GPUs (8–24 GB), use batch size 1 with gradient accumulation to simulate larger effective batches without extra VRAM.

Does gradient checkpointing affect model quality?

No. Gradient checkpointing produces mathematically identical gradients and model weights. It trades ~20–30% slower training for several-fold less activation memory by recomputing activations during the backward pass. It does not shrink the logits, which dominate at long context unless the framework chunks them (Unsloth does). Always recommended.

Can I use system RAM for fine-tuning instead of VRAM?

Generally no. CPU offloading exists via DeepSpeed ZeRO-Offload, but it is 10–100× slower than GPU training. Fine-tuning requires high-bandwidth memory access that only GPU VRAM (or unified memory on Apple Silicon) provides. System RAM is useful for dataset loading, not the training loop.

Can I fine-tune LLMs on a Mac with Apple Silicon?

Yes, with MLX. Unsloth and bitsandbytes need NVIDIA. macOS lets the GPU use about 75% of unified memory (≈13.5 GB of 18 GB), and sequence length matters as much as model size: an 18 GB Mac trains an 8B model with QLoRA only up to a few thousand tokens. Pick a Mac preset above to see each model's limit. Training is 2–4× slower than NVIDIA GPUs.

Can I train a 70B model on an RTX 4090 (24GB)?

No. Even with QLoRA and Unsloth, a 70B model needs about 42.0 GB at 4,096-token sequences — an RTX 4090 has 24 GB. Use an A100/H100 80GB or a Mac with 96 GB+ unified memory (MLX). On 24 GB the largest dense models that fit at 4,096 tokens are about 33B with QLoRA (QwQ 32B) and 11B with LoRA 16-bit (Llama 3.2 Vision 11B).

What are PEFT and LoRA? How do they reduce memory?

PEFT (Parameter-Efficient Fine-Tuning) is an umbrella term for methods that train only a small subset of parameters. LoRA (Low-Rank Adaptation) injects small trainable matrices into each transformer layer while freezing the original weights. This reduces trainable parameters by ~95–99%, cutting optimizer memory proportionally. QLoRA adds 4-bit quantization on top for even more savings.

Sources & Methodology

VRAM and fine-tuning estimates on this page are based on peer-reviewed research and established frameworks. Key references:

Estimates are approximations based on standard implementations. Actual VRAM usage varies by framework (PyTorch, JAX, MLX), version, and hardware. Always profile your specific workload.

Related Tools

Related Guides