Every memory figure, fit verdict and speed ceiling on this site, as files you can download, query and cite. The architecture values of the 29 models are read from each model's own config.json on Hugging Face, the memory and bandwidth of the 20 GPUs from the manufacturers' specification pages, and everything else is computed from those with the formula below. Nothing is benchmarked, and nothing is copied from another site's table. Free to use under CC BY 4.0, commercial use included, with credit to Nodegrove.
Archived on Zenodo, DOI 10.5281/zenodo.22966137. The DOI always points at the newest version.
What the data shows
Newer models pay far less for context. At 32k tokens, Qwen3 32B (2025, standard attention) spends 8.6 GB on its KV cache; its successor Qwen3.8 27B (2026, hybrid attention) spends 2.3 GB. Across all 29 models, each 1,000 tokens of context costs between 0.006 GB (Nemotron 3.5 Lightning 30B-A3B) and 0.328 GB (DeepSeek-R1 Distill Llama 70B).
Active parameters set the speed, not the memory. Qwen3-Coder-Next 80B-A3B reads 3B parameters per token but keeps all 79.7B in memory: 48.9 GB at Q4_K_M with 8k context, against 22.4 GB for the dense Qwen3 32B.
24 of 29 models fit a 24 GB card at Q4_K_M with 8k context. The largest is Qwen3.6 35B-A3B at 22.4 GB, under the 22.8 GB line (95% of the card).
Files
- llm-vram-models.csv29 rows
One row per model: size, attention design, context window, licence and the source of every value.
- llm-vram-gpus.csv20 rows
One row per GPU: memory, the memory a runtime can use, bandwidth and the specification page.
- llm-vram-requirements.csv1,068 rows
Memory per model, quantisation and context length, split into weights, KV cache and overhead.
- llm-vram-gpu-fit.csv3,480 rows
Every model on every GPU at every quantisation, at 8k context: verdict, longest context that fits and speed ceiling.
-
Everything above in one file, with the method, its constants and each model's cache layout.
This is version 2026-09-25. Each file is served from nodegrove.io with open CORS, so a notebook or an app can read it by URL, and every version is kept as a dated release on GitHub. The data is refreshed when new open models come out.
Read it from code
Which models run on an RTX 4090 at Q4_K_M, in Python:
import pandas as pd
fit = pd.read_csv("https://nodegrove.io/data/llm-vram-gpu-fit.csv")
fit.query("gpu_id == 'rtx-4090' and quant == 'Q4_K_M' and verdict != 'no'") Or the whole dataset in JavaScript:
const url = 'https://nodegrove.io/data/llm-vram.json';
const data = await fetch(url).then((r) => r.json()); How the numbers are made
Memory is weights plus KV cache plus overhead, the same arithmetic as the VRAM calculator.
- Weights are parameters multiplied by bytes per parameter: FP16 2.00, Q8_0 1.06, Q6_K 0.82, Q5_K_M 0.71, Q4_K_M 0.58, Q3_K_M 0.47. These are effective averages for GGUF files, including the scales and the unquantised embedding and output layers.
- KV cache is, for each group of layers, the values cached per token multiplied by the tokens held and by 2 bytes. 14 of the 29 models are standard transformers, where every layer caches keys and values for the whole context. The other 15 cache less and are counted as they cache: sliding-window layers stop at their window (7 models), hybrid models keep a growing cache on their full-attention layers only, plus a small fixed state (6), and latent attention caches one compressed vector per token (2).
- Overhead is 0.5 GB for the runtime plus 4% of the weights for activations and buffers.
- Fit. A model fits a GPU when its total at 8k context is at most 95% of the memory the runtime can address, and is tight above 85%. Apple silicon gives the GPU about 75% of its unified memory by default, and the dataset counts that share.
- Speed ceiling is 0.7 × memory bandwidth ÷ bytes of active weights, for one stream generating text. It is a ceiling, not a measurement: the faster the figure, the further real runtimes fall below it. The speed estimator explains why.
These are estimates from published values. Real usage moves with the runtime, the batch size, flash attention and cache quantisation, so read "fits" as "worth trying". The savings of sliding-window, hybrid and latent attention also depend on the runtime implementing the matching cache: current llama.cpp and vLLM do for these families, and an older build may store every layer at full length. Why newer models pay less for context has the working.
Columns
The JSON holds the same rows under models, gpus, requirements and gpu_fit, plus the method, its constants and each model's cache layout (kv_groups). Blank means not applicable.
llm-vram-models.csv
- model_id
- Stable identifier, also the page slug on nodegrove.io.
- model
- Model name.
- family
- Model family.
- released
- Month the weights were published (YYYY-MM).
- params_b
- Total parameters, billions, as the checkpoint reports them. Mixture-of-experts models count every expert; multimodal checkpoints include the vision encoder.
- active_params_b
- Parameters read per token, billions. Mixture-of-experts models only; blank for dense models.
- attention
- How the model caches context: standard, sliding-window, hybrid or latent (see the method).
- layers
- Transformer layers (blocks).
- kv_heads
- Key/value heads of the main attention layers.
- head_dim
- Dimension of each key/value head.
- full_cache_layers
- Layers whose cache grows with the whole context.
- window_layers
- Layers that only keep the last window_tokens tokens.
- window_tokens
- Sliding-window length in tokens; blank if the model has none.
- kv_cache_gb_per_1k_tokens
- Memory each extra 1,000 tokens of context costs, GB, with an FP16 cache once any windows are full.
- fixed_state_gb
- Fixed recurrent state of linear-attention or Mamba layers, GB. Does not grow with context.
- context_tokens
- Native context window in tokens (max_position_embeddings, or the model card where the two differ).
- license
- Licence of the weights, from the model card.
- license_permissive
- true when the licence has no field-of-use or scale conditions.
- role
- code or reasoning for specialist models; blank for general-purpose models.
- successor_id
- Newer model in the same family and role, if there is one.
- hf_repo
- Hugging Face repository the values were read from.
- hf_gated
- true when config.json opens only after accepting the licence on Hugging Face.
- config_url
- Link to the config.json the architecture values come from.
- page_url
- The model page on nodegrove.io.
- attention_detail
- The attention layout in one sentence.
- notes
- Caveats that change how the numbers should be read.
llm-vram-gpus.csv
- gpu_id
- Stable identifier, also the page slug on nodegrove.io.
- gpu
- Card or machine.
- kind
- consumer, workstation, datacenter, apple, or unified (non-Apple unified memory).
- maker
- NVIDIA, AMD, Apple or Intel.
- memory_gb
- Memory from the manufacturer specification, GB (unified memory for Apple silicon).
- usable_gb
- Memory an inference runtime can address, GB. Apple silicon: about 75% of unified memory (observed; Apple publishes no figure). Other unified-memory machines: the share their maker documents, stated on the GPU page.
- budget_gb
- 95% of usable_gb: the line a model must fit under to count as fitting.
- bandwidth_gb_s
- Memory bandwidth from the manufacturer specification, GB/s.
- spec_url
- The manufacturer's specification page.
- bandwidth_source_url
- Where the bandwidth figure is printed, when spec_url does not print it (a maker whitepaper, datasheet or announcement, or the data-rate source for a stated bus × rate calculation). Empty when spec_url prints it.
- page_url
- The GPU page on nodegrove.io.
- notes
- Caveats.
llm-vram-requirements.csv
- model_id
- Joins llm-vram-models.csv.
- model
- Model name.
- quant
- FP16, Q8_0, Q6_K, Q5_K_M, Q4_K_M or Q3_K_M.
- bytes_per_param
- Effective bytes per parameter at that quantisation.
- context_tokens
- Tokens held in context.
- weights_gb
- Parameters × bytes_per_param, GB.
- kv_cache_gb
- KV cache stored at FP16, plus any fixed state, GB.
- overhead_gb
- Runtime context and buffers: 0.5 GB + 4% of weights.
- total_gb
- weights_gb + kv_cache_gb + overhead_gb.
- total_gb_q8_cache
- The same total with the KV cache quantised to 8 bits.
llm-vram-gpu-fit.csv
- model_id
- Joins llm-vram-models.csv.
- model
- Model name.
- gpu_id
- Joins llm-vram-gpus.csv.
- gpu
- Card or machine.
- quant
- Quantisation of the weights.
- context_tokens
- Context the fit is computed at (8,192 tokens).
- need_gb
- Estimated memory at that quantisation and context, GB (total_gb in llm-vram-requirements.csv).
- verdict
- fits: need_gb is within the GPU's budget_gb. tight: it fits, but above 85% of usable memory. no: it does not fit.
- max_context_tokens
- Longest context that fits at this quantisation with an FP16 cache, capped at 131,072 tokens or the model's window. 0 when the weights alone do not fit.
- tokens_per_second_ceiling
- Upper bound on single-stream generation speed, tokens/s (see the method). Blank when the model cannot load.
Licence and citation
Use, share and adapt it under CC BY 4.0, commercially too, as long as you credit Nodegrove and link to this page. To cite it:
Nodegrove (2026). LLM VRAM dataset (version 2026-09-25) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.22966137 Each model's weights carry their own licence, listed per row; the dataset contains no model files. Model and GPU names are trademarks of their owners. Found a value that disagrees with its source? Open an issue or write to info@nodegrove.io with the row and the link, and the next version will carry the fix.
- Model architecture: each model's config.json and model card on Hugging Face, linked per row in config_url.
- GPU memory and bandwidth: the manufacturers' specification pages, linked per row in spec_url.
- Formulas: the same code as the VRAM calculator, the can-I-run-it checker and the speed estimator.