The LLM VRAM dataset

Memory for 29 open models at six quantisations and every common context length, and which of 20 GPUs run each one. CSV and JSON, CC BY 4.0.

Updated · estimates are labelled as estimates
On this page

Every memory figure, fit verdict and speed ceiling on this site, as files you can download, query and cite. The architecture values of the 29 models are read from each model's own config.json on Hugging Face, the memory and bandwidth of the 20 GPUs from the manufacturers' specification pages, and everything else is computed from those with the formula below. Nothing is benchmarked, and nothing is copied from another site's table. Free to use under CC BY 4.0, commercial use included, with credit to Nodegrove.

Archived on Zenodo, DOI 10.5281/zenodo.22966137. The DOI always points at the newest version.

What the data shows

Newer models pay far less for context. At 32k tokens, Qwen3 32B (2025, standard attention) spends 8.6 GB on its KV cache; its successor Qwen3.8 27B (2026, hybrid attention) spends 2.3 GB. Across all 29 models, each 1,000 tokens of context costs between 0.006 GB (Nemotron 3.5 Lightning 30B-A3B) and 0.328 GB (DeepSeek-R1 Distill Llama 70B).

Active parameters set the speed, not the memory. Qwen3-Coder-Next 80B-A3B reads 3B parameters per token but keeps all 79.7B in memory: 48.9 GB at Q4_K_M with 8k context, against 22.4 GB for the dense Qwen3 32B.

24 of 29 models fit a 24 GB card at Q4_K_M with 8k context. The largest is Qwen3.6 35B-A3B at 22.4 GB, under the 22.8 GB line (95% of the card).

Files

  • One row per model: size, attention design, context window, licence and the source of every value.

  • One row per GPU: memory, the memory a runtime can use, bandwidth and the specification page.

  • Memory per model, quantisation and context length, split into weights, KV cache and overhead.

  • Every model on every GPU at every quantisation, at 8k context: verdict, longest context that fits and speed ceiling.

  • Everything above in one file, with the method, its constants and each model's cache layout.

This is version 2026-09-25. Each file is served from nodegrove.io with open CORS, so a notebook or an app can read it by URL, and every version is kept as a dated release on GitHub. The data is refreshed when new open models come out.

Read it from code

Which models run on an RTX 4090 at Q4_K_M, in Python:

import pandas as pd

fit = pd.read_csv("https://nodegrove.io/data/llm-vram-gpu-fit.csv")
fit.query("gpu_id == 'rtx-4090' and quant == 'Q4_K_M' and verdict != 'no'")

Or the whole dataset in JavaScript:

const url = 'https://nodegrove.io/data/llm-vram.json';
const data = await fetch(url).then((r) => r.json());

How the numbers are made

Memory is weights plus KV cache plus overhead, the same arithmetic as the VRAM calculator.

  • Weights are parameters multiplied by bytes per parameter: FP16 2.00, Q8_0 1.06, Q6_K 0.82, Q5_K_M 0.71, Q4_K_M 0.58, Q3_K_M 0.47. These are effective averages for GGUF files, including the scales and the unquantised embedding and output layers.
  • KV cache is, for each group of layers, the values cached per token multiplied by the tokens held and by 2 bytes. 14 of the 29 models are standard transformers, where every layer caches keys and values for the whole context. The other 15 cache less and are counted as they cache: sliding-window layers stop at their window (7 models), hybrid models keep a growing cache on their full-attention layers only, plus a small fixed state (6), and latent attention caches one compressed vector per token (2).
  • Overhead is 0.5 GB for the runtime plus 4% of the weights for activations and buffers.
  • Fit. A model fits a GPU when its total at 8k context is at most 95% of the memory the runtime can address, and is tight above 85%. Apple silicon gives the GPU about 75% of its unified memory by default, and the dataset counts that share.
  • Speed ceiling is 0.7 × memory bandwidth ÷ bytes of active weights, for one stream generating text. It is a ceiling, not a measurement: the faster the figure, the further real runtimes fall below it. The speed estimator explains why.

These are estimates from published values. Real usage moves with the runtime, the batch size, flash attention and cache quantisation, so read "fits" as "worth trying". The savings of sliding-window, hybrid and latent attention also depend on the runtime implementing the matching cache: current llama.cpp and vLLM do for these families, and an older build may store every layer at full length. Why newer models pay less for context has the working.

Columns

The JSON holds the same rows under models, gpus, requirements and gpu_fit, plus the method, its constants and each model's cache layout (kv_groups). Blank means not applicable.

llm-vram-models.csv

model_id
Stable identifier, also the page slug on nodegrove.io.
model
Model name.
family
Model family.
released
Month the weights were published (YYYY-MM).
params_b
Total parameters, billions, as the checkpoint reports them. Mixture-of-experts models count every expert; multimodal checkpoints include the vision encoder.
active_params_b
Parameters read per token, billions. Mixture-of-experts models only; blank for dense models.
attention
How the model caches context: standard, sliding-window, hybrid or latent (see the method).
layers
Transformer layers (blocks).
kv_heads
Key/value heads of the main attention layers.
head_dim
Dimension of each key/value head.
full_cache_layers
Layers whose cache grows with the whole context.
window_layers
Layers that only keep the last window_tokens tokens.
window_tokens
Sliding-window length in tokens; blank if the model has none.
kv_cache_gb_per_1k_tokens
Memory each extra 1,000 tokens of context costs, GB, with an FP16 cache once any windows are full.
fixed_state_gb
Fixed recurrent state of linear-attention or Mamba layers, GB. Does not grow with context.
context_tokens
Native context window in tokens (max_position_embeddings, or the model card where the two differ).
license
Licence of the weights, from the model card.
license_permissive
true when the licence has no field-of-use or scale conditions.
role
code or reasoning for specialist models; blank for general-purpose models.
successor_id
Newer model in the same family and role, if there is one.
hf_repo
Hugging Face repository the values were read from.
hf_gated
true when config.json opens only after accepting the licence on Hugging Face.
config_url
Link to the config.json the architecture values come from.
page_url
The model page on nodegrove.io.
attention_detail
The attention layout in one sentence.
notes
Caveats that change how the numbers should be read.

llm-vram-gpus.csv

gpu_id
Stable identifier, also the page slug on nodegrove.io.
gpu
Card or machine.
kind
consumer, workstation, datacenter, apple, or unified (non-Apple unified memory).
maker
NVIDIA, AMD, Apple or Intel.
memory_gb
Memory from the manufacturer specification, GB (unified memory for Apple silicon).
usable_gb
Memory an inference runtime can address, GB. Apple silicon: about 75% of unified memory (observed; Apple publishes no figure). Other unified-memory machines: the share their maker documents, stated on the GPU page.
budget_gb
95% of usable_gb: the line a model must fit under to count as fitting.
bandwidth_gb_s
Memory bandwidth from the manufacturer specification, GB/s.
spec_url
The manufacturer's specification page.
bandwidth_source_url
Where the bandwidth figure is printed, when spec_url does not print it (a maker whitepaper, datasheet or announcement, or the data-rate source for a stated bus × rate calculation). Empty when spec_url prints it.
page_url
The GPU page on nodegrove.io.
notes
Caveats.

llm-vram-requirements.csv

model_id
Joins llm-vram-models.csv.
model
Model name.
quant
FP16, Q8_0, Q6_K, Q5_K_M, Q4_K_M or Q3_K_M.
bytes_per_param
Effective bytes per parameter at that quantisation.
context_tokens
Tokens held in context.
weights_gb
Parameters × bytes_per_param, GB.
kv_cache_gb
KV cache stored at FP16, plus any fixed state, GB.
overhead_gb
Runtime context and buffers: 0.5 GB + 4% of weights.
total_gb
weights_gb + kv_cache_gb + overhead_gb.
total_gb_q8_cache
The same total with the KV cache quantised to 8 bits.

llm-vram-gpu-fit.csv

model_id
Joins llm-vram-models.csv.
model
Model name.
gpu_id
Joins llm-vram-gpus.csv.
gpu
Card or machine.
quant
Quantisation of the weights.
context_tokens
Context the fit is computed at (8,192 tokens).
need_gb
Estimated memory at that quantisation and context, GB (total_gb in llm-vram-requirements.csv).
verdict
fits: need_gb is within the GPU's budget_gb. tight: it fits, but above 85% of usable memory. no: it does not fit.
max_context_tokens
Longest context that fits at this quantisation with an FP16 cache, capped at 131,072 tokens or the model's window. 0 when the weights alone do not fit.
tokens_per_second_ceiling
Upper bound on single-stream generation speed, tokens/s (see the method). Blank when the model cannot load.

Licence and citation

Use, share and adapt it under CC BY 4.0, commercially too, as long as you credit Nodegrove and link to this page. To cite it:

Nodegrove (2026). LLM VRAM dataset (version 2026-09-25) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.22966137

Each model's weights carry their own licence, listed per row; the dataset contains no model files. Model and GPU names are trademarks of their owners. Found a value that disagrees with its source? Open an issue or write to info@nodegrove.io with the row and the link, and the next version will carry the fix.

  • Model architecture: each model's config.json and model card on Hugging Face, linked per row in config_url.
  • GPU memory and bandwidth: the manufacturers' specification pages, linked per row in spec_url.
  • Formulas: the same code as the VRAM calculator, the can-I-run-it checker and the speed estimator.