Already know the model and the card? The can-I-run-it checker answers that pairing directly. Below, all 29 open models tracked on this site are measured the same way, at Q4_K_M with 8k of context, from Llama 3.2 3B at 3.4 GB to Mistral Small 4 119B (MoE) at 72.5 GB. 24 of them fit a 24 GB card at that setting. Nothing here is benchmarked: memory comes from a stated formula over each model's published architecture, and every model's page links the file the numbers were read from.
Every model, compared
| Model | Size | Q4 · 8k | Q8 · 8k | Per 1k context | Window | Smallest card |
|---|---|---|---|---|---|---|
| Llama 3.2 3B | 3.2B | 3.4 GB | 5 GB | 0.115 GB | 128k | 8 GB |
| Qwen3.5 4B | 4.7B | 3.7 GB | 6 GB | 0.033 GB | 256k | 8 GB |
| Llama 3.1 8B | 8B | 6.4 GB | 10.4 GB | 0.131 GB | 128k | 8 GB |
| Qwen3 8B | 8.2B | 6.7 GB | 10.7 GB | 0.147 GB | 40k | 8 GB |
| Qwen3.5 9B | 9.7B | 6.7 GB | 11.5 GB | 0.033 GB | 256k | 8 GB |
| Gemma 4 12B | 12B | 8.2 GB | 14.2 GB | 0.016 GB | 256k | 12 GB |
| Gemma 3 12B | 12.2B | 8.7 GB | 14.8 GB | 0.066 GB | 128k | 12 GB |
| Ministral 3 14B | 14B | 10.3 GB | 17.3 GB | 0.164 GB | 256k | 12 GB |
| Qwen3 14B | 14.8B | 10.8 GB | 18.2 GB | 0.164 GB | 40k | 12 GB |
| Phi-4 14B | 14.7B | 11 GB | 18.4 GB | 0.205 GB | 16k | 12 GB |
| gpt-oss 20B | 21B · 3.6B active | 13.4 GB | 23.9 GB | 0.025 GB | 128k | 16 GB |
| Devstral Small 2 24B | 24B | 16.3 GB | 28.3 GB | 0.164 GB | 256k | 24 GB |
| Mistral Small 3.1 24B | 24B | 16.3 GB | 28.3 GB | 0.164 GB | 128k | 24 GB |
| Gemma 4 26B-A4B (MoE) | 25.8B · 3.8B active | 16.4 GB | 29.3 GB | 0.02 GB | 256k | 24 GB |
| Qwen3.8 27B | 27.8B | 18 GB | 31.8 GB | 0.066 GB | 256k | 24 GB |
| Gemma 3 27B | 27.4B | 18.1 GB | 31.8 GB | 0.082 GB | 128k | 24 GB |
| Qwen3 30B-A3B (MoE) | 30.5B · 3.3B active | 19.7 GB | 34.9 GB | 0.098 GB | 40k | 24 GB |
| Nemotron 3.5 Lightning 30B-A3B (MoE) | 31.6B · 3B active | 19.7 GB | 35.4 GB | 0.006 GB | 256k | 24 GB |
| GLM-4.7 Flash 30B-A3B (MoE) | 31.2B · 3B active | 19.8 GB | 35.3 GB | 0.054 GB | 198k | 24 GB |
| Gemma 4 31B | 31.3B | 20.9 GB | 36.5 GB | 0.082 GB | 256k | 24 GB |
| Qwen3 32B | 32.8B | 22.4 GB | 38.8 GB | 0.262 GB | 40k | 24 GB |
| Qwen2.5 Coder 32B | 32.8B | 22.4 GB | 38.8 GB | 0.262 GB | 32k | 24 GB |
| DeepSeek-R1 Distill Qwen 32B | 32.8B | 22.4 GB | 38.8 GB | 0.262 GB | 128k | 24 GB |
| Qwen3.6 35B-A3B (MoE) | 36B · 3B active | 22.4 GB | 40.4 GB | 0.02 GB | 256k | 24 GB |
| Llama 3.3 70B | 70.6B | 45.8 GB | 81 GB | 0.328 GB | 128k | 80 GB |
| DeepSeek-R1 Distill Llama 70B | 70.6B | 45.8 GB | 81 GB | 0.328 GB | 128k | 80 GB |
| Qwen3-Coder-Next 80B-A3B (MoE) | 79.7B · 3B active | 48.9 GB | 88.6 GB | 0.025 GB | 256k | 80 GB |
| gpt-oss 120B | 117B · 5.1B active | 71.4 GB | 129.8 GB | 0.037 GB | 128k | 80 GB |
| Mistral Small 4 119B (MoE) | 119B · 6.5B active | 72.5 GB | 131.9 GB | 0.023 GB | 256k | 80 GB |
"Smallest card" is the smallest memory size that holds the model at Q4_K_M with 8k of context. For named cards rather than sizes, see every GPU compared; for 32k context and the full fit table, see which models fit on 8 to 96 GB.
How to read the table
The Q4 column decides whether it loads. Weights are the big number: parameters multiplied by the bytes each takes at the chosen quantisation. A mixture-of-experts model keeps every expert in memory, so its full size counts here even though only the active share is read per token.
The context column decides how long a conversation it holds. It is the memory each extra 1,000 tokens costs, and it is where models of the same size differ most: Nemotron 3.5 Lightning 30B-A3B (MoE) pays 0.006 GB per 1,000 tokens, Llama 3.3 70B pays 0.328 GB. Models released in 2026 mostly keep a growing cache on only some of their layers, which is why they sit at the cheap end. Why newer models pay less for context has the working.
The window is a limit of the model, not of your GPU. Spare memory does not extend it. No page on this site credits a model with more context than its card states.
Qwen
- Qwen3.8 27B 18 GB at Q4 · fits 24 GB · 256k window · Apache 2.0 · 2026-08
Alibaba's dense 27B from August 2026, and one of the most downloaded open models of the year.
- Qwen3.6 35B-A3B (MoE) 22.4 GB at Q4 · fits 24 GB · 256k window · Apache 2.0 · 2026-04
The first open-weight Qwen3.6 model, April 2026: 256 experts with 8 routed and one shared, 3B parameters read per token.
- Qwen3.5 4B 3.7 GB at Q4 · fits 8 GB · 256k window · Apache 2.0 · 2026-02
The small end of Alibaba's Qwen3.5 family from February 2026.
- Qwen3.5 9B 6.7 GB at Q4 · fits 8 GB · 256k window · Apache 2.0 · 2026-02
The 9B model of Alibaba's Qwen3.5 family, February 2026, and one of the most downloaded open models of the year.
- Qwen3-Coder-Next 80B-A3B (MoE) 48.9 GB at Q4 · fits 80 GB · 256k window · Apache 2.0 · 2026-01 · code model
Alibaba's coding-agent model, early 2026: 80B of weights across 512 experts, with 3B read per token.
- Qwen3 8B 6.7 GB at Q4 · fits 8 GB · 40k window · Apache 2.0 · 2025-04
The 8B dense model of Alibaba's Qwen3 generation, April 2025.
Newer in this family: Qwen3.5 9B - Qwen3 14B 10.8 GB at Q4 · fits 12 GB · 40k window · Apache 2.0 · 2025-04
The 14B dense model of the Qwen3 generation, April 2025.
- Qwen3 30B-A3B (MoE) 19.7 GB at Q4 · fits 24 GB · 40k window · Apache 2.0 · 2025-04
The small mixture-of-experts model of the Qwen3 generation, April 2025: 30B of weights, 3.3B read per token.
Newer in this family: Qwen3.6 35B-A3B (MoE) - Qwen3 32B 22.4 GB at Q4 · fits 24 GB · 40k window · Apache 2.0 · 2025-04
The dense flagship of the Qwen3 generation, April 2025, and for a year the model a 24 GB card was bought to run.
Newer in this family: Qwen3.8 27B - Qwen2.5 Coder 32B 22.4 GB at Q4 · fits 24 GB · 32k window · Apache 2.0 · 2024-11 · code model
Alibaba's coding model from November 2024, and the open coding model of choice through most of 2025.
Newer in this family: Qwen3-Coder-Next 80B-A3B (MoE)
Gemma
- Gemma 4 12B 8.2 GB at Q4 · fits 12 GB · 256k window · Apache 2.0 · 2026-05
Google DeepMind's Gemma 4 12B "Unified", May 2026.
- Gemma 4 26B-A4B (MoE) 16.4 GB at Q4 · fits 24 GB · 256k window · Apache 2.0 · 2026-03
The mixture-of-experts member of Gemma 4, spring 2026: 128 experts with 8 active plus one shared, 3.8B parameters read per token.
- Gemma 4 31B 20.9 GB at Q4 · fits 24 GB · 256k window · Apache 2.0 · 2026-03
The dense flagship of Google DeepMind's Gemma 4 family, spring 2026.
- Gemma 3 12B 8.7 GB at Q4 · fits 12 GB · 128k window · Gemma Terms of Use · 2025-03
Google's 12B vision-language model from March 2025.
Newer in this family: Gemma 4 12B - Gemma 3 27B 18.1 GB at Q4 · fits 24 GB · 128k window · Gemma Terms of Use · 2025-03
Google's largest Gemma 3 model, March 2025: vision-language, 128k window, and for a year the most common answer to "what do I run on a 24 GB card".
Newer in this family: Gemma 4 31B
Mistral
- Mistral Small 4 119B (MoE) 72.5 GB at Q4 · fits 80 GB · 256k window · Apache 2.0 · 2026-03
Mistral's March 2026 general model, and small in name only: 119B of weights, 128 experts, 6.5B read per token.
- Ministral 3 14B 10.3 GB at Q4 · fits 12 GB · 256k window · Apache 2.0 · 2025-12
The largest of Mistral's Ministral 3 family, December 2025, designed for edge and local deployment.
- Devstral Small 2 24B 16.3 GB at Q4 · fits 24 GB · 256k window · Apache 2.0 · 2025-12 · code model
Mistral's agentic coding model, December 2025.
- Mistral Small 3.1 24B 16.3 GB at Q4 · fits 24 GB · 128k window · Apache 2.0 · 2025-03
Mistral's 24B general model from March 2025: text and images, a 128k window, Apache 2.0.
Newer in this family: Mistral Small 4 119B (MoE)
Llama
- Llama 3.3 70B 45.8 GB at Q4 · fits 80 GB · 128k window · Llama Community License · 2024-12
Meta's 70B text model from December 2024, and still the reference dense 70B. When people ask whether a card "can run a 70B", this is the model they mean.
- Llama 3.2 3B 3.4 GB at Q4 · fits 8 GB · 128k window · Llama Community License · 2024-09
Meta's small text model from September 2024, built to run on phones and laptops.
- Llama 3.1 8B 6.4 GB at Q4 · fits 8 GB · 128k window · Llama Community License · 2024-07
Meta's 8B text model from July 2024, and probably the most fine-tuned open model there has been.
DeepSeek
- DeepSeek-R1 Distill Qwen 32B 22.4 GB at Q4 · fits 24 GB · 128k window · MIT · 2025-01 · reasoning model
DeepSeek-R1's reasoning distilled into Qwen2.5 32B, January 2025.
- DeepSeek-R1 Distill Llama 70B 45.8 GB at Q4 · fits 80 GB · 128k window · MIT · 2025-01 · reasoning model
DeepSeek-R1's reasoning distilled into Llama 3.3 70B, January 2025.
gpt-oss
- gpt-oss 20B 13.4 GB at Q4 · fits 16 GB · 128k window · Apache 2.0 · 2025-08
The smaller of OpenAI's two open-weight models, August 2025.
- gpt-oss 120B 71.4 GB at Q4 · fits 80 GB · 128k window · Apache 2.0 · 2025-08
The larger of OpenAI's open-weight models, August 2025: 117B of weights, 5.1B read per token, shipped in MXFP4 so that it fits a single 80 GB card.
GLM
- GLM-4.7 Flash 30B-A3B (MoE) 19.8 GB at Q4 · fits 24 GB · 198k window · MIT · 2026-01
Z.ai's 30B-A3B mixture-of-experts model, January 2026, the lightweight member of the GLM-4.7 line.
Nemotron
- Nemotron 3.5 Lightning 30B-A3B (MoE) 19.7 GB at Q4 · fits 24 GB · 256k window · OpenMDW 1.1 · 2026-08
NVIDIA's open 30B-A3B model, August 2026, released with its training data and recipes.
Phi
- Phi-4 14B 11 GB at Q4 · fits 12 GB · 16k window · MIT · 2024-12
Microsoft's 14B model from December 2024, trained largely on synthetic data with an emphasis on reasoning and mathematics.
How to choose, in one paragraph
Start from the memory you have, not the model you have heard of. Find the largest row whose Q4 figure fits your card with room to spare, then check its context column against how long your documents and conversations really are. If two models fit, the newer one usually holds more conversation in the same memory. If the model you want does not fit anything you own, the checker lists exactly what would change that, and the build vs rent calculator tells you honestly whether buying a bigger card beats renting one for the hours you would use it.
- Architecture values are read from each model's config.json and model card on Hugging Face; every model page links its own source.
- Memory: weights + KV cache + runtime overhead, per the formula in the VRAM calculator. Sliding-window, hybrid and latent-attention models are counted as they actually cache.
- Context cost is at an FP16 KV cache, once any sliding windows are full.
- Every figure in this table, and each model at every quantisation and context length, is also open data: CSV and JSON under CC BY 4.0, archived on Zenodo with a DOI.