Can I Fine-Tune LLM is a free VRAM estimator for QLoRA, LoRA, and full fine-tuning — 95 models, RTX 5090/4090 and Apple Silicon, 100% in your browser.
Estimate GPU memory for QLoRA, LoRA & Full
Fine-Tuning of open-source LLMs
Configuration
Apple Silicon: macOS lets the GPU use about 75% of unified memory —
13.5 GB of 18 GB. Unsloth and
bitsandbytes need NVIDIA, so estimates use MLX (mlx_lm.lora).
GB
2GB192GB
64GB
4GB256GB
4,096tokens
512128K
1
132
COMFORTABLE
TIGHT
OOM
Top Picks for Your Hardware
Training Compatibility — All Models
Quick Reference: VRAM Needed for Local
Fine-Tuning
Estimated GPU memory for common model sizes. Unsloth, batch size 1,
4K-token sequences, gradient checkpointing on.
Model Size
QLoRA
(4-bit)
LoRA
(8-bit)
LoRA
(16-bit)
Full
Fine-Tune
Min GPU
(QLoRA)
1B
~2.5 GB
~2.8 GB
~3.5 GB
~16.4 GB
RTX 4060 (8GB)
7B
~7.8 GB
~10.7 GB
~16.8 GB
~115.6 GB
RTX 3060 (12GB)
8B
~7.7 GB
~10.8 GB
~17.3 GB
~121.1 GB
RTX 3060 (12GB)
14B
~12.1 GB
~18.1 GB
~30.4 GB
~222.2 GB
RTX 4060 Ti (16GB)
32B
~21.8 GB
~35.5 GB
~63.9 GB
~479.4 GB
RTX 4090 (24GB)
70B
~42.0 GB
~72.6 GB
~135.8 GB
~1,046 GB
RTX 6000 Ada (48GB)
Unsloth, LoRA rank 16 on all linear layers, batch 1, 4,096-token
sequences, gradient checkpointing on. Rows use real architectures (1B = Llama 3.2 1B, 7B = Qwen 2.5 7B, 8B = Llama 3.1 8B, 14B = Qwen 3 14B, 32B = Qwen 2.5 32B, 70B = Llama 3.3 70B). Plain Hugging Face +
PEFT needs more — mostly for the full-vocabulary logits — and memory grows with every extra token of
context: use the calculator above for your exact setup. Estimates are fitted to Unsloth's published
context-length benchmarks (8B and 70B, ±10%); small models under Hugging Face/MLX are likely overestimated.
How we calculate?
Fine-tuning VRAM = Model Weights + Adapter Optimizer States + Activations × sequence length +
Logits + Framework Overhead, computed from each model's real architecture (its config.json) and
fitted to Unsloth's published context-length benchmarks.
QLoRA quantizes base weights to 4-bit (NF4) and only trains small LoRA adapter matrices.
LoRA keeps the base frozen at 16-bit and trains adapters. Full fine-tuning updates all weights,
requiring AdamW optimizer states (8 bytes/param) and full gradients (4 bytes/param).
For MoE models, ALL expert weights are trained — not just the active fraction.
What is Gradient Checkpointing?
Instead of storing all intermediate activations during the forward pass, gradient checkpointing
keeps each layer's input and recomputes the rest during the backward pass. This saves most of the
activation memory (not the logits) at the cost of
~20–30% slower training. The output is mathematically identical — no quality loss. Always
recommended for local training on consumer GPUs.
QLoRA vs LoRA vs Full FT
QLoRA (4-bit): Best for consumer GPUs. ~0.5 bytes/param base. Trains 1–5% of parameters
via adapters. LoRA (8-bit): Middle ground. ~1 byte/param base. Slightly better quality than
QLoRA. LoRA (16-bit): bf16 base, 2 bytes/param. Best adapter quality. Full FT: Updates all weights. ~16 bytes/param. Enterprise-only.
All methods add activation memory per token of context — see the calculator.
QLoRA vs LoRA vs Full Fine-Tuning: VRAM Requirements Compared
The fine-tuning method you choose has a 5–10× impact on VRAM consumption — which directly determines
whether you can run a given model on consumer hardware at all. QLoRA, LoRA, and full fine-tuning occupy
three very different points on the memory–quality tradeoff curve. Understanding where each sits lets you
make the right call before you spin up a training run and hit an OOM error two hours in.
Method
Base Model Precision
Adapter Precision
VRAM for 7B Model
VRAM for 13B
VRAM for 70B
Notes
Full Fine-Tuning
FP32
N/A
~112 GB
~208 GB
~1.12 TB
Needs multi-GPU for anything >3B
Full Fine-Tuning
BF16
N/A
~56 GB
~104 GB
~560 GB
Still impractical without enterprise GPU
LoRA (FP16 base)
FP16
FP32 adapters
~16–18 GB
~28–32 GB
~140–160 GB
Popular for 24 GB consumer GPUs
QLoRA (4-bit base)
NF4 4-bit
BF16 adapters
~6–8 GB
~12–14 GB
~48–56 GB
Consumer-grade fine-tuning
QLoRA + GC
NF4 4-bit + GC
BF16 adapters
~5–7 GB
~10–12 GB
~40–48 GB
GC = gradient checkpointing
QLoRA achieves its dramatic memory savings through NF4 (NormalFloat4) quantization applied to the frozen
base model weights. NF4 is not the same as INT4 — it uses a quantization grid that is information-theoretically
optimal for normally distributed weights, which is what you find in pretrained LLMs. Because the base model
is frozen and only the small LoRA adapter matrices are updated, the adapter gradients are kept in full
BF16 precision to preserve training stability. The quantization error introduced by NF4 is partially
compensated by double quantization, where the quantization constants themselves are quantized a second
time, saving an additional 0.37 bits per parameter.
LoRA without quantization is worth the extra VRAM in two situations. First, when convergence speed
matters — FP16 base models typically converge 10–20% faster than QLoRA because the optimizer sees
full-precision gradients throughout. Second, when final model quality is the priority: on tasks that
demand very precise domain adaptation (e.g., legal or medical text), LoRA consistently scores 1–3
percentage points higher than QLoRA on downstream benchmarks. If you have a 24 GB GPU and are running
a 7B model, LoRA fits comfortably and the quality gains are usually worth it. For anything larger,
QLoRA is the only viable path without multi-GPU setup.
Gradient Checkpointing: Trading Training Speed for VRAM Savings
During a standard forward pass, the model stores every intermediate activation tensor in VRAM so the
backward pass can use them to compute gradients. For a 7B model with a typical batch size and sequence
length, this activation memory can add 5–15 GB on top of the model weights themselves — easily pushing
a workload over the edge on a 16 GB or 24 GB card. Gradient checkpointing (GC) solves this by discarding
most intermediate activations as they are computed and recomputing them on-the-fly during the backward
pass. The result is a 30–60% reduction in peak VRAM usage from activations, which can be the difference
between fitting a model on consumer hardware and not fitting it at all.
The cost of this VRAM saving is training time. Because activations must be recomputed, gradient
checkpointing typically adds 20–40% to total training time compared to a run without it. In practice,
the recommendation is straightforward: always enable gradient checkpointing when fine-tuning on consumer
hardware (RTX 4090 and below, or Apple M-series with 24 GB or less). The time penalty is acceptable
when it unlocks the ability to train at all. Disable gradient checkpointing only when VRAM is genuinely
plentiful — for example, on an A100 80 GB or H100 with a model that already fits with headroom to spare
and where throughput is more important than memory frugality.
Enabling gradient checkpointing in Hugging Face Trainer is a single flag. The
use_reentrant=False setting is strongly recommended when using QLoRA because the reentrant
autograd checkpoint implementation has known incompatibilities with the bitsandbytes quantization hooks
that QLoRA relies on:
The gradient_accumulation_steps=4 setting above is a complementary technique: it simulates
a larger effective batch size (4 in this case) without storing more activations simultaneously, because
only one micro-batch is live in memory at any given moment. Combining gradient checkpointing with
gradient accumulation is the standard playbook for fine-tuning large models on consumer hardware, and
both are supported natively by TRL's SFTTrainer without any additional code.
Fine-Tuning on Apple Silicon: MLX vs MPS (PyTorch)
Apple M-series chips use a unified memory architecture where the GPU and CPU share the same physical RAM
pool. This is an advantage for LLM fine-tuning, with one catch: by default macOS lets the GPU use only about
75% of it (Metal's working-set limit) — about 13.5 GB on an 18 GB Mac or 36 GB on a 48 GB M4 Max — and the
OS and your apps need the rest. The limit can be raised with sudo sysctl iogpu.wired_limit_mb,
at the risk of starving macOS. Even so, a 48 GB Mac gives a far bigger "GPU" than any consumer NVIDIA card.
Chip
Unified Memory
Max Model for QLoRA (short sequences)
MLX Support
MPS (PyTorch) Support
Practical Throughput
M3 / M4 (base)
16 GB
7B Q4
✓
✓ (limited)
~150–200 tok/s training
M3 Pro / M4 Pro
18–24 GB
7B–13B Q4
✓
✓
~200–350 tok/s
M3 Max / M4 Max
36–48 GB
13B–34B Q4
✓
✓
~400–700 tok/s
M2 Ultra / M3 Ultra
64–192 GB
Up to 70B Q4
✓
✓
~800–1500 tok/s
MLX is the preferred framework for fine-tuning on Apple Silicon over PyTorch's MPS (Metal Performance
Shaders) backend for several reasons. MLX is written from the ground up for Apple's Metal GPU API and
uses native Metal ops throughout the entire compute graph. PyTorch's MPS backend, by contrast, was
originally designed for CUDA and uses a translation layer that introduces overhead, particularly for
non-standard ops that fall back to CPU. In practice, MLX achieves 20–40% higher token throughput than
PyTorch MPS on the same Apple Silicon hardware, and it avoids the MPS memory fragmentation bugs that
can cause out-of-memory errors mid-run even when total memory use is within budget.
Fine-tuning a model with MLX requires just two steps after installing the package. Install with
pip install mlx-lm, then launch a LoRA fine-tuning run with a single command:
mlx_lm.lora --model mistralai/Mistral-7B-v0.1 --train --data ./data/. MLX will
automatically download the model from Hugging Face, apply LoRA adapters, and train using the JSONL
files in your ./data/ directory (which should contain train.jsonl and
optionally valid.jsonl). Adapter weights are saved as adapters.npz and can
be merged back into the base model with mlx_lm.fuse for inference.
MLX does have meaningful limitations to plan around. It does not support multi-GPU or multi-node training
— all compute happens on the single unified memory pool of one machine. As of the MLX 0.x series, Flash
Attention is not implemented, which means memory usage for long sequences scales quadratically rather
than linearly and very long context fine-tuning (32K+ tokens) is not practical. Not all model
architectures are supported yet; Llama, Mistral, Phi, Qwen, and Gemma families work well, but newer or
less popular architectures may require a custom conversion. For these cases, falling back to PyTorch MPS
or a cloud GPU is the practical path.
Local vs Cloud Fine-Tuning: Cost Breakdown for 2026
For a single fine-tuning run on a 7B model with 10,000 training examples — roughly 3–4 hours of
wall-clock training time on a single high-end consumer GPU — the economics vary dramatically depending
on whether you own your hardware or rent cloud compute. The table below compares the most common
options available in 2026, with prices in both INR and USD for reference.
Option
Setup Cost
Cost per Run
Best For
Caveats
RTX 4090 (owned)
₹1.5–2L one-time
₹0 (electricity only, ~₹20–40)
Frequent fine-tuning
High upfront cost
MacBook M4 Max 48 GB
₹3.5–4.5L one-time
₹0 (electricity, ~₹10–20)
Mac-only workflows
Slower than CUDA for LLMs
RunPod (RTX 4090)
₹0
~₹120–200/run ($1.5–2.5/hr × 3–4 hr)
Occasional runs
Cold start, data upload
Vast.ai (RTX 3090)
₹0
~₹60–100/run ($0.7–1.2/hr × 3–4 hr)
Budget runs
Less reliable uptime
Lambda Labs (A100 80 GB)
₹0
~₹500–700/run ($2–2.5/hr × 3–4 hr)
Large models >34B
Enterprise grade
Google Colab Pro+
~₹1700/month
Shared with subscription
Students, hobbyists
Session time limits, slow disk
The break-even calculation for owning versus renting depends primarily on how many fine-tuning runs you
plan to do per year. At RunPod RTX 4090 prices of ~₹150 per run, an RTX 4090 card costing ₹1.8L
amortizes after roughly 50–70 runs — which is less than one run per week over a year. If you are
actively iterating on models (multiple runs per week for hyperparameter search, dataset experiments, or
production model updates), owning the hardware pays off within a few months. If you run fine-tuning
only a handful of times per quarter, renting cloud GPUs is almost always cheaper when you factor in
electricity, cooling, and the opportunity cost of capital tied up in hardware.
Data privacy is the other major factor that the raw cost numbers do not capture. Cloud fine-tuning
platforms — RunPod, Vast.ai, Lambda Labs, and Colab — all require you to upload your training dataset
to their infrastructure. For general-purpose or open-domain fine-tuning this is usually acceptable.
But if your training data contains proprietary business information, customer data, medical records,
legal documents, or anything regulated under GDPR, HIPAA, or similar frameworks, uploading it to a
third-party cloud platform creates compliance and legal exposure that may outweigh the cost savings
entirely. In those cases, local fine-tuning — whether on an owned GPU or Apple Silicon — is not just
a preference but a requirement. The RTX 4090 or M4 Max starts to look very cost-effective the moment
data sovereignty becomes a hard constraint.
Dataset Preparation: The Step That Makes or Breaks Fine-Tuning
Even with perfect hardware and the right fine-tuning method, the output quality of your fine-tuned model
is bounded by the quality of your training data. This is the part of the pipeline that practitioners
underestimate most often. A model fine-tuned on noisy, inconsistent, or poorly formatted data will
perform worse than the base model on your target task — a phenomenon sometimes called "fine-tuning
degradation." Getting dataset preparation right is therefore not optional; it is the core engineering
challenge of supervised fine-tuning.
Format: Instruction-Tuning vs Completion
The two dominant data formats for fine-tuning are instruction-response pairs and completion-only
sequences. Instruction-response (also called chat or instruction-tuning format) structures each example
as a system prompt plus a user message plus an expected model response. This format aligns with how
modern chat models are trained and is the right choice when you want to teach the model to follow
instructions in a specific style or domain. Completion-only format is simpler: it is a continuous block
of text that the model learns to continue, and it is appropriate for tasks like code generation or
domain-specific text continuation where instruction-following behaviour is already present in the base
model and you just want to shift the output distribution.
Volume and Quality Thresholds
The minimum viable dataset size for meaningful fine-tuning is roughly 500–1000 high-quality examples
for a narrow, well-defined task (e.g., reformatting product descriptions into a specific JSON schema).
For broader behavioural changes — teaching a model a new domain, persona, or reasoning style — you
typically need 5,000–50,000 examples. The quality threshold matters far more than quantity: 500 carefully
curated, consistent, human-verified examples routinely outperform 10,000 examples scraped from the web
without filtering. A practical rule of thumb is to hand-review at least 5% of your dataset before
training — if you find errors in more than 5% of reviewed examples, the whole dataset likely needs
another cleaning pass.
Deduplication and Train/Validation Split
Duplicate examples in a fine-tuning dataset are more damaging than in large-scale pretraining because
the training corpus is much smaller and the model is more susceptible to memorising repeated patterns.
Always deduplicate at the level of both exact matches and near-duplicates (using MinHash or a similar
approximate deduplication method). Reserve 5–10% of your data as a held-out validation set and monitor
validation loss during training — if training loss keeps falling but validation loss starts rising, the
model is overfitting to your training examples and you should stop early. A train/validation split
ratio of 90/10 is standard; using a completely separate test set for final evaluation (not touched
during training or hyperparameter selection) is best practice for production deployments.
Data Format for MLX and Hugging Face
Both MLX and Hugging Face TRL expect training data as JSONL files where each line is a JSON object.
For instruction-tuning with Hugging Face TRL's SFTTrainer, the standard format is a
messages field containing a list of role/content dictionaries, matching the OpenAI chat
format. For MLX, each line should have a text field with the fully formatted prompt
already applied. Converting between formats is straightforward, but applying the correct chat template
(using the tokenizer's apply_chat_template method) is essential — using the wrong
template or formatting prompts manually without the template often introduces invisible whitespace or
special token differences that silently degrade model quality at inference time.
Sequence Length and Packing
Maximum sequence length is one of the most impactful hyperparameters for both VRAM usage and training
throughput. Setting it too high wastes VRAM on padding tokens and reduces the number of examples that
fit per batch; setting it too low truncates examples and loses information. A practical approach is to
compute the 95th-percentile token length across your dataset and set max_seq_length to
that value, then review the 5% of examples that are truncated to decide if they are critical. For most
instruction-tuning datasets with short to medium-length responses, a value of 1024–2048 tokens strikes
the right balance between VRAM efficiency and data completeness. If your dataset has many short
examples, enabling sequence packing — concatenating multiple short examples into a single training
sequence separated by EOS tokens — can dramatically improve GPU utilisation and reduce training time
by 30–50% without affecting model quality.
LLM Fine-Tuning FAQ & VRAM Guide
What is the difference between QLoRA and LoRA
VRAM usage?
QLoRA stores the frozen base in 4-bit NF4 (~0.5 bytes/param) and trains small
LoRA adapters. LoRA 16-bit keeps the base in bf16 (2 bytes/param).
Full fine-tuning needs ~16 bytes/param (weights, gradients, fp32 master copy and
Adam states). On top of that, every method needs activation memory that grows with sequence
length — often the part that decides whether a run fits.
How much VRAM to fine-tune Llama 3 8B?
At 4,096-token sequences with gradient checkpointing: QLoRA needs about 7.7 GB with Unsloth or 19.4 GB with plain Hugging Face + PEFT; LoRA 16-bit about 17.3 GB (Unsloth); full fine-tuning about 121.1 GB. Context length matters as much as model size: on a 24 GB GPU, QLoRA fits roughly 81K tokens with Unsloth but only about 5.9K with plain Hugging Face.
Does batch size affect VRAM during
fine-tuning?
Yes, significantly. Activation memory scales linearly with batch size. Doubling
the batch size roughly doubles activation memory. For consumer GPUs (8–24 GB), use batch size 1
with gradient accumulation to simulate larger effective batches without extra
VRAM.
Does gradient checkpointing affect model
quality?
No. Gradient checkpointing produces mathematically identical
gradients and model weights. It trades ~20–30% slower training for several-fold less activation
memory by recomputing activations during the backward pass. It does not shrink the logits, which
dominate at long context unless the framework chunks them (Unsloth does). Always recommended.
Can I use system RAM for fine-tuning instead
of VRAM?
Generally no. CPU offloading exists via DeepSpeed ZeRO-Offload, but it is
10–100× slower than GPU training. Fine-tuning requires high-bandwidth memory access that only
GPU VRAM (or unified memory on Apple Silicon) provides. System RAM is useful for dataset
loading, not the training loop.
Can I fine-tune LLMs on a Mac with Apple
Silicon?
Yes, with MLX. Unsloth and bitsandbytes need NVIDIA. macOS lets the GPU use
about 75% of unified memory (≈13.5 GB of 18 GB), and sequence length matters as much as model
size: an 18 GB Mac trains an 8B model with QLoRA only up to a few thousand tokens. Pick a Mac
preset above to see each model's limit. Training is 2–4× slower than NVIDIA GPUs.
Can I train a 70B model on an RTX 4090
(24GB)?
No. Even with QLoRA and Unsloth, a 70B model needs about 42.0 GB at 4,096-token sequences — an RTX 4090 has 24 GB. Use an A100/H100 80GB or a Mac with 96 GB+ unified memory (MLX). On 24 GB the largest dense models that fit at 4,096 tokens are about 33B with QLoRA (QwQ 32B) and 11B with LoRA 16-bit (Llama 3.2 Vision 11B).
What are PEFT and LoRA? How do they reduce
memory?
PEFT (Parameter-Efficient Fine-Tuning) is an umbrella term for methods that
train only a small subset of parameters. LoRA (Low-Rank Adaptation) injects
small trainable matrices into each transformer layer while freezing the original weights. This
reduces trainable parameters by ~95–99%, cutting optimizer memory proportionally. QLoRA adds
4-bit quantization on top for even more savings.
Sources & Methodology
VRAM and fine-tuning estimates on this page are based on peer-reviewed research and established frameworks. Key references:
Estimates are approximations based on standard implementations. Actual VRAM usage varies by framework (PyTorch, JAX, MLX), version, and hardware. Always profile your specific workload.