How much VRAM does each open model need?

29 open models measured the same way: memory at Q4 and Q8, what each 1,000 tokens of context costs, and the smallest GPU that runs each one.

Updated · estimates are labelled as estimates
On this page

Already know the model and the card? The can-I-run-it checker answers that pairing directly. Below, all 29 open models tracked on this site are measured the same way, at Q4_K_M with 8k of context, from Llama 3.2 3B at 3.4 GB to Mistral Small 4 119B (MoE) at 72.5 GB. 24 of them fit a 24 GB card at that setting. Nothing here is benchmarked: memory comes from a stated formula over each model's published architecture, and every model's page links the file the numbers were read from.

Every model, compared

ModelSizeQ4 · 8kQ8 · 8kPer 1k contextWindowSmallest card
Llama 3.2 3B 3.2B 3.4 GB 5 GB 0.115 GB 128k 8 GB
Qwen3.5 4B 4.7B 3.7 GB 6 GB 0.033 GB 256k 8 GB
Llama 3.1 8B 8B 6.4 GB 10.4 GB 0.131 GB 128k 8 GB
Qwen3 8B 8.2B 6.7 GB 10.7 GB 0.147 GB 40k 8 GB
Qwen3.5 9B 9.7B 6.7 GB 11.5 GB 0.033 GB 256k 8 GB
Gemma 4 12B 12B 8.2 GB 14.2 GB 0.016 GB 256k 12 GB
Gemma 3 12B 12.2B 8.7 GB 14.8 GB 0.066 GB 128k 12 GB
Ministral 3 14B 14B 10.3 GB 17.3 GB 0.164 GB 256k 12 GB
Qwen3 14B 14.8B 10.8 GB 18.2 GB 0.164 GB 40k 12 GB
Phi-4 14B 14.7B 11 GB 18.4 GB 0.205 GB 16k 12 GB
gpt-oss 20B 21B · 3.6B active 13.4 GB 23.9 GB 0.025 GB 128k 16 GB
Devstral Small 2 24B 24B 16.3 GB 28.3 GB 0.164 GB 256k 24 GB
Mistral Small 3.1 24B 24B 16.3 GB 28.3 GB 0.164 GB 128k 24 GB
Gemma 4 26B-A4B (MoE) 25.8B · 3.8B active 16.4 GB 29.3 GB 0.02 GB 256k 24 GB
Qwen3.8 27B 27.8B 18 GB 31.8 GB 0.066 GB 256k 24 GB
Gemma 3 27B 27.4B 18.1 GB 31.8 GB 0.082 GB 128k 24 GB
Qwen3 30B-A3B (MoE) 30.5B · 3.3B active 19.7 GB 34.9 GB 0.098 GB 40k 24 GB
Nemotron 3.5 Lightning 30B-A3B (MoE) 31.6B · 3B active 19.7 GB 35.4 GB 0.006 GB 256k 24 GB
GLM-4.7 Flash 30B-A3B (MoE) 31.2B · 3B active 19.8 GB 35.3 GB 0.054 GB 198k 24 GB
Gemma 4 31B 31.3B 20.9 GB 36.5 GB 0.082 GB 256k 24 GB
Qwen3 32B 32.8B 22.4 GB 38.8 GB 0.262 GB 40k 24 GB
Qwen2.5 Coder 32B 32.8B 22.4 GB 38.8 GB 0.262 GB 32k 24 GB
DeepSeek-R1 Distill Qwen 32B 32.8B 22.4 GB 38.8 GB 0.262 GB 128k 24 GB
Qwen3.6 35B-A3B (MoE) 36B · 3B active 22.4 GB 40.4 GB 0.02 GB 256k 24 GB
Llama 3.3 70B 70.6B 45.8 GB 81 GB 0.328 GB 128k 80 GB
DeepSeek-R1 Distill Llama 70B 70.6B 45.8 GB 81 GB 0.328 GB 128k 80 GB
Qwen3-Coder-Next 80B-A3B (MoE) 79.7B · 3B active 48.9 GB 88.6 GB 0.025 GB 256k 80 GB
gpt-oss 120B 117B · 5.1B active 71.4 GB 129.8 GB 0.037 GB 128k 80 GB
Mistral Small 4 119B (MoE) 119B · 6.5B active 72.5 GB 131.9 GB 0.023 GB 256k 80 GB

"Smallest card" is the smallest memory size that holds the model at Q4_K_M with 8k of context. For named cards rather than sizes, see every GPU compared; for 32k context and the full fit table, see which models fit on 8 to 96 GB.

How to read the table

The Q4 column decides whether it loads. Weights are the big number: parameters multiplied by the bytes each takes at the chosen quantisation. A mixture-of-experts model keeps every expert in memory, so its full size counts here even though only the active share is read per token.

The context column decides how long a conversation it holds. It is the memory each extra 1,000 tokens costs, and it is where models of the same size differ most: Nemotron 3.5 Lightning 30B-A3B (MoE) pays 0.006 GB per 1,000 tokens, Llama 3.3 70B pays 0.328 GB. Models released in 2026 mostly keep a growing cache on only some of their layers, which is why they sit at the cheap end. Why newer models pay less for context has the working.

The window is a limit of the model, not of your GPU. Spare memory does not extend it. No page on this site credits a model with more context than its card states.

Qwen

  • Qwen3.8 27B
    18 GB at Q4 · fits 24 GB · 256k window · Apache 2.0 · 2026-08

    Alibaba's dense 27B from August 2026, and one of the most downloaded open models of the year.

  • Qwen3.6 35B-A3B (MoE)
    22.4 GB at Q4 · fits 24 GB · 256k window · Apache 2.0 · 2026-04

    The first open-weight Qwen3.6 model, April 2026: 256 experts with 8 routed and one shared, 3B parameters read per token.

  • Qwen3.5 4B
    3.7 GB at Q4 · fits 8 GB · 256k window · Apache 2.0 · 2026-02

    The small end of Alibaba's Qwen3.5 family from February 2026.

  • Qwen3.5 9B
    6.7 GB at Q4 · fits 8 GB · 256k window · Apache 2.0 · 2026-02

    The 9B model of Alibaba's Qwen3.5 family, February 2026, and one of the most downloaded open models of the year.

  • Qwen3-Coder-Next 80B-A3B (MoE)
    48.9 GB at Q4 · fits 80 GB · 256k window · Apache 2.0 · 2026-01 · code model

    Alibaba's coding-agent model, early 2026: 80B of weights across 512 experts, with 3B read per token.

  • Qwen3 8B
    6.7 GB at Q4 · fits 8 GB · 40k window · Apache 2.0 · 2025-04

    The 8B dense model of Alibaba's Qwen3 generation, April 2025.

    Newer in this family: Qwen3.5 9B
  • Qwen3 14B
    10.8 GB at Q4 · fits 12 GB · 40k window · Apache 2.0 · 2025-04

    The 14B dense model of the Qwen3 generation, April 2025.

  • Qwen3 30B-A3B (MoE)
    19.7 GB at Q4 · fits 24 GB · 40k window · Apache 2.0 · 2025-04

    The small mixture-of-experts model of the Qwen3 generation, April 2025: 30B of weights, 3.3B read per token.

    Newer in this family: Qwen3.6 35B-A3B (MoE)
  • Qwen3 32B
    22.4 GB at Q4 · fits 24 GB · 40k window · Apache 2.0 · 2025-04

    The dense flagship of the Qwen3 generation, April 2025, and for a year the model a 24 GB card was bought to run.

    Newer in this family: Qwen3.8 27B
  • Qwen2.5 Coder 32B
    22.4 GB at Q4 · fits 24 GB · 32k window · Apache 2.0 · 2024-11 · code model

    Alibaba's coding model from November 2024, and the open coding model of choice through most of 2025.

    Newer in this family: Qwen3-Coder-Next 80B-A3B (MoE)

Gemma

  • Gemma 4 12B
    8.2 GB at Q4 · fits 12 GB · 256k window · Apache 2.0 · 2026-05

    Google DeepMind's Gemma 4 12B "Unified", May 2026.

  • Gemma 4 26B-A4B (MoE)
    16.4 GB at Q4 · fits 24 GB · 256k window · Apache 2.0 · 2026-03

    The mixture-of-experts member of Gemma 4, spring 2026: 128 experts with 8 active plus one shared, 3.8B parameters read per token.

  • Gemma 4 31B
    20.9 GB at Q4 · fits 24 GB · 256k window · Apache 2.0 · 2026-03

    The dense flagship of Google DeepMind's Gemma 4 family, spring 2026.

  • Gemma 3 12B
    8.7 GB at Q4 · fits 12 GB · 128k window · Gemma Terms of Use · 2025-03

    Google's 12B vision-language model from March 2025.

    Newer in this family: Gemma 4 12B
  • Gemma 3 27B
    18.1 GB at Q4 · fits 24 GB · 128k window · Gemma Terms of Use · 2025-03

    Google's largest Gemma 3 model, March 2025: vision-language, 128k window, and for a year the most common answer to "what do I run on a 24 GB card".

    Newer in this family: Gemma 4 31B

Mistral

  • Mistral Small 4 119B (MoE)
    72.5 GB at Q4 · fits 80 GB · 256k window · Apache 2.0 · 2026-03

    Mistral's March 2026 general model, and small in name only: 119B of weights, 128 experts, 6.5B read per token.

  • Ministral 3 14B
    10.3 GB at Q4 · fits 12 GB · 256k window · Apache 2.0 · 2025-12

    The largest of Mistral's Ministral 3 family, December 2025, designed for edge and local deployment.

  • Devstral Small 2 24B
    16.3 GB at Q4 · fits 24 GB · 256k window · Apache 2.0 · 2025-12 · code model

    Mistral's agentic coding model, December 2025.

  • Mistral Small 3.1 24B
    16.3 GB at Q4 · fits 24 GB · 128k window · Apache 2.0 · 2025-03

    Mistral's 24B general model from March 2025: text and images, a 128k window, Apache 2.0.

    Newer in this family: Mistral Small 4 119B (MoE)

Llama

  • Llama 3.3 70B
    45.8 GB at Q4 · fits 80 GB · 128k window · Llama Community License · 2024-12

    Meta's 70B text model from December 2024, and still the reference dense 70B. When people ask whether a card "can run a 70B", this is the model they mean.

  • Llama 3.2 3B
    3.4 GB at Q4 · fits 8 GB · 128k window · Llama Community License · 2024-09

    Meta's small text model from September 2024, built to run on phones and laptops.

  • Llama 3.1 8B
    6.4 GB at Q4 · fits 8 GB · 128k window · Llama Community License · 2024-07

    Meta's 8B text model from July 2024, and probably the most fine-tuned open model there has been.

DeepSeek

  • DeepSeek-R1 Distill Qwen 32B
    22.4 GB at Q4 · fits 24 GB · 128k window · MIT · 2025-01 · reasoning model

    DeepSeek-R1's reasoning distilled into Qwen2.5 32B, January 2025.

  • DeepSeek-R1 Distill Llama 70B
    45.8 GB at Q4 · fits 80 GB · 128k window · MIT · 2025-01 · reasoning model

    DeepSeek-R1's reasoning distilled into Llama 3.3 70B, January 2025.

gpt-oss

  • gpt-oss 20B
    13.4 GB at Q4 · fits 16 GB · 128k window · Apache 2.0 · 2025-08

    The smaller of OpenAI's two open-weight models, August 2025.

  • gpt-oss 120B
    71.4 GB at Q4 · fits 80 GB · 128k window · Apache 2.0 · 2025-08

    The larger of OpenAI's open-weight models, August 2025: 117B of weights, 5.1B read per token, shipped in MXFP4 so that it fits a single 80 GB card.

GLM

  • GLM-4.7 Flash 30B-A3B (MoE)
    19.8 GB at Q4 · fits 24 GB · 198k window · MIT · 2026-01

    Z.ai's 30B-A3B mixture-of-experts model, January 2026, the lightweight member of the GLM-4.7 line.

Nemotron

  • Nemotron 3.5 Lightning 30B-A3B (MoE)
    19.7 GB at Q4 · fits 24 GB · 256k window · OpenMDW 1.1 · 2026-08

    NVIDIA's open 30B-A3B model, August 2026, released with its training data and recipes.

Phi

  • Phi-4 14B
    11 GB at Q4 · fits 12 GB · 16k window · MIT · 2024-12

    Microsoft's 14B model from December 2024, trained largely on synthetic data with an emphasis on reasoning and mathematics.

How to choose, in one paragraph

Start from the memory you have, not the model you have heard of. Find the largest row whose Q4 figure fits your card with room to spare, then check its context column against how long your documents and conversations really are. If two models fit, the newer one usually holds more conversation in the same memory. If the model you want does not fit anything you own, the checker lists exactly what would change that, and the build vs rent calculator tells you honestly whether buying a bigger card beats renting one for the hours you would use it.

  • Architecture values are read from each model's config.json and model card on Hugging Face; every model page links its own source.
  • Memory: weights + KV cache + runtime overhead, per the formula in the VRAM calculator. Sliding-window, hybrid and latent-attention models are counted as they actually cache.
  • Context cost is at an FP16 KV cache, once any sliding windows are full.
  • Every figure in this table, and each model at every quantisation and context length, is also open data: CSV and JSON under CC BY 4.0, archived on Zenodo with a DOI.