Mistral Small 4 119B (MoE) has 119B parameters with 6.5B active per token (mixture of experts). It has 36 layers using multi-head latent attention: instead of keys and values for each of its 32 heads, every layer caches one compressed vector of 320 values per token. At the everyday setting, Q4_K_M and 8k context, it needs an estimated 72.5 GB, so the smallest card that fits is 80 GB.
What Mistral Small 4 119B (MoE) is
Mistral's March 2026 general model, and small in name only: 119B of weights, 128 experts, 6.5B read per token. It folds three earlier lines into one model: Instruct, the Magistral reasoning models, and Devstral for code.
Text and images 262,144-token window Apache 2.0 March 2026 6.5B of 119B active
Reasoning. Switchable per request between instant replies and a reasoning mode with configurable effort.
Memory by quantisation and context
| Quant | 4,096 ctx | 8,192 ctx | 32,768 ctx |
|---|---|---|---|
| Q8_0 | 131.8 GB · fits multi-GPU | 131.9 GB · fits multi-GPU | 132.4 GB · fits multi-GPU |
| Q5_K_M | 88.5 GB · fits 96 GB | 88.6 GB · fits 96 GB | 89.1 GB · fits 96 GB |
| Q4_K_M | 72.4 GB · fits 80 GB | 72.5 GB · fits 80 GB | 73 GB · fits 80 GB |
Weights at this quant: 69 GB at Q4. Every extra 1,000 tokens of context adds about 0.023 GB of KV cache at FP16. Try other settings in the VRAM calculator.
Which GPU to pick for Mistral Small 4 119B (MoE)
The smallest card here that loads it at Q4_K_M with 8k of context is the A100 (80 GB): 72.5 GB needed, an estimated 379 tokens per second, and a tight fit that leaves little for the conversation. For room to work, meaning 32,768 tokens of context while staying under 85% of memory, step up to the RTX PRO 6000 Blackwell (96 GB), which needs 73 GB for that and generates at about 333 tokens per second.
At Q8_0 it does not fit any single card listed here. On Apple silicon, a MacBook Pro M4 Max holds it with 32,768 tokens of context at an estimated 101 tokens per second; prompt processing is slower than on Nvidia, which long documents make obvious. Among unified-memory desktops, the Ryzen AI Max+ 395 holds it at 32,768 tokens in 73 GB, at an estimated 48 tokens per second. To check a card that is not on this page, or a different context length, use the can-I-run-it checker, which opens on this model.
Speed by GPU, at Q4 and 8k context
| GPU | Memory | Bandwidth | Fits | Est. tokens/s | Feels like |
|---|---|---|---|---|---|
| RTX 3060 12 GB | 12 GB | 360 GB/s | no | – | does not fit |
| RTX 4060 Ti 16 GB | 16 GB | 288 GB/s | no | – | does not fit |
| RTX 4070 12 GB | 12 GB | 504 GB/s | no | – | does not fit |
| RTX 3090 24 GB | 24 GB | 936 GB/s | no | – | does not fit |
| RTX 4090 24 GB | 24 GB | 1,008 GB/s | no | – | does not fit |
| RTX 5090 32 GB | 32 GB | 1,792 GB/s | no | – | does not fit |
| RTX 6000 Ada 48 GB | 48 GB | 960 GB/s | no | – | does not fit |
| L40S 48 GB | 48 GB | 864 GB/s | no | – | does not fit |
| RTX PRO 6000 Blackwell 96 GB | 96 GB | 1,792 GB/s | yes | ≤ 333 | faster than you can read |
| A100 80 GB | 80 GB | 2,039 GB/s | tight | ≤ 379 | faster than you can read |
| H100 SXM 80 GB | 80 GB | 3,350 GB/s | tight | ≤ 622 | faster than you can read |
| RTX 5060 Ti 16 GB | 16 GB | 448 GB/s | no | – | does not fit |
| Radeon RX 7900 XTX 24 GB | 24 GB | 960 GB/s | no | – | does not fit |
| DGX Spark 128 GB (unified) | 128 GB (126 for the GPU) | 273 GB/s | yes | 51 | comfortable for chat |
| Ryzen AI Max+ 395 128 GB (unified) | 128 GB (96 for the GPU) | 256 GB/s | yes | 48 | comfortable for chat |
Single-stream decode, from memory bandwidth at 70% efficiency. Prompt processing and batching not included. These are ceilings, not measurements: Mistral Small 4 119B (MoE) reads only 3.8 GB of weights per token at Q4, so little that the per-token costs the formula leaves out decide much of the real speed. The 3 figures marked ≤ are where a real runtime falls furthest below the number shown. See the speed estimator for other quantisations and Apple hardware, or every card compared if you are choosing hardware rather than a model.
What Mistral Small 4 119B (MoE) is good at, and what it is not
Good at
- One model instead of three on an 80 GB card: chat, reasoning and coding.
- Long context almost for free: latent attention caches one small vector per layer.
- A 128 GB Mac, for the same reason as every large mixture-of-experts model.
- Apache 2.0, and Mistral publishes an NVFP4 checkpoint and a speculative-decoding head for it.
Watch out for
- The name invites a mistake. Its predecessor ran on a 24 GB card; this needs 80 GB.
- The latent cache needs runtime support. Without it, context memory is many times higher.
- Its config allows a million positions; Mistral's card states 256k, and that is the figure used here.
Similar models
The models that need about the same memory as Mistral Small 4 119B (MoE) at Q4_K_M and 8k, which makes them the real alternatives on whatever card you have. The context column is what each extra 1,000 tokens costs, and it is where models of the same size differ most.
| Model | Size | Needs | Per 1k context | Window | Released |
|---|---|---|---|---|---|
| Mistral Small 4 119B (MoE) | 119B · 6.5B active | 72.5 GB | 0.023 GB | 256k | 2026-03 |
| gpt-oss 120B | 117B · 5.1B active | 71.4 GB | 0.037 GB | 128k | 2025-08 |
| Qwen3-Coder-Next 80B-A3B (MoE) | 79.7B · 3B active | 48.9 GB | 0.025 GB | 256k | 2026-01 |
| DeepSeek-R1 Distill Llama 70B | 70.6B | 45.8 GB | 0.328 GB | 128k | 2025-01 |
| Llama 3.3 70B | 70.6B | 45.8 GB | 0.328 GB | 128k | 2024-12 |
Notes
- Small in name only: all 128 experts stay in memory. A runtime that does not implement the latent cache needs several times more context memory. The parameter count includes the vision encoder; a text-only GGUF is slightly smaller.
- Native context window: 262,144 tokens. Licence: Apache 2.0, a permissive licence.
- Weights published in March 2026. It follows Mistral Small 3.1 24B.
- Architecture values from config.json in mistralai/Mistral-Small-4-119B-2603 on Hugging Face, and its model card. Every figure on this page is computed from those values; check them before buying hardware for this model.
- Estimates assume a single conversation on a card that is otherwise free. A desktop on the same GPU takes 0.5 to 2 GB.
Questions
Can I run Mistral Small 4 119B (MoE) on a 24 GB card?
Not comfortably. At Q4_K_M and 8k context it needs about 72.5 GB. The smallest tier that fits is 80 GB.
How much VRAM does Mistral Small 4 119B (MoE) need at Q8?
About 131.9 GB at 8k context, or 132.4 GB at 32k. Q8 is near-lossless; use it when it fits.
What is the context length of Mistral Small 4 119B (MoE)?
262,144 tokens natively. Holding all of it at Q4_K_M takes about 78.3 GB, of which 6 GB is context.
Why does Mistral Small 4 119B (MoE) need so little memory for long context?
Because most of its layers do not keep a cache that grows. It has 36 layers using multi-head latent attention: instead of keys and values for each of its 32 heads, every layer caches one compressed vector of 320 values per token. At 32k tokens the context costs about 0.8 GB; if all 36 layers kept a conventional full-length cache, the same conversation would cost about 19.3 GB.
Can I run Mistral Small 4 on a 24 GB GPU?
No. Despite the name it is a 119B-parameter model, and every expert has to sit in memory even though only 6.5B parameters are used per token. At Q4 it needs an 80 GB card or a large-memory Mac. On a 24 GB card, Devstral Small 2 and Mistral Small 3.1 are the Mistral models that fit.
How fast is Mistral Small 4 119B (MoE) on an RTX 4090?
It does not fit a 4090 at Q4 and 8k context, so speed would collapse to CPU offloading. Use a larger card or a smaller quantisation.
See every model compared and which models fit on 8 to 96 GB, or run it without buying the card: Nodegrove attaches a 24, 48 or 96 GB GPU to a workspace that stays saved.