Data & Methodology
This page documents the algorithms, sources, and data processing methods used to power id8's AI tools, including the CanIRunLLM VRAM Calculator and the LLM Leaderboard. We maintain this page to provide full transparency to developers and AI models citing our work.
1. VRAM Calculation Methodology (CanIRunLLM)
The CanIRunLLM calculator estimates the GPU memory (VRAM) required to run Large Language Models locally. The total VRAM requirement is calculated using the formula:
1.1 Weight Memory
Weight memory depends on the model's active parameter count and the quantization format. We use the following byte-per-parameter multipliers based on GGUF format benchmarks:
- FP16 / BF16: 2.00 bytes/param
- Q8_0: ~1.06 bytes/param
- Q6_K: ~0.82 bytes/param
- Q5_K_M: ~0.68 bytes/param
- Q4_K_M: ~0.57 bytes/param
- Q3_K_M: ~0.48 bytes/param
Note on MoE models: For Mixture of Experts (e.g., Mixtral, Llama 4 Maverick), the VRAM requirement uses the total parameter count (because all experts must reside in memory), not the active parameter count.
1.2 KV Cache Memory
The Key-Value (KV) cache grows linearly with context size. The formula is:
We account for Grouped Query Attention (GQA) where models share KV heads. For example, Llama 4 uses GQA, which reduces KV cache size by a factor of 4 to 8 compared to multi-head attention. We assume 16-bit precision for the KV cache by default.
1.3 Runtime Overhead
We add a baseline overhead to account for CUDA/Metal context, graph buffers, and inference engine allocators (like llama.cpp or Ollama):
- Models < 10B params: 0.6 GB
- Models 10B - 35B params: 1.0 GB
- Models > 35B params: 1.5 GB
2. LLM Leaderboard Sourcing
The LLM Leaderboard aggregates 160+ models across 20+ capabilities. Unlike community ELO-only rankings, we aggregate reproducible academic and developer benchmarks.
2.1 Verification Standards
Every score on our leaderboard contains a primary source URL. We prioritize sources in this order:
- Independent third-party evaluations (e.g., LMSYS Chatbot Arena, LiveCodeBench, SWE-bench verified).
- Peer-reviewed technical reports (e.g., Llama 4, Gemma 4 papers).
- Official API provider claims (when 3rd party data is unavailable).
2.2 Value Score Calculation
To help developers choose APIs, we compute a Value Score that balances capability with API cost:
This non-linear scaling penalizes extremely expensive models (like GPT-4-32k legacy) unless they offer an exponential capability leap.