Ministral 3 14B has 14B parameters, 40 layers and 8 KV heads of dimension 128. At the everyday setting, Q4_K_M and 8k context, it needs an estimated 10.3 GB, so the smallest card that fits is 12 GB.
What Ministral 3 14B is
The largest of Mistral's Ministral 3 family, December 2025, designed for edge and local deployment. A conventional dense transformer with vision, a 256k window and an Apache licence.
Text and images 262,144-token window Apache 2.0 December 2025
Reasoning. The Instruct release answers directly. Mistral publishes a separate Reasoning variant of the same size.
Memory by quantisation and context
| Quant | 4,096 ctx | 8,192 ctx | 32,768 ctx |
|---|---|---|---|
| Q8_0 | 16.6 GB · fits 24 GB | 17.3 GB · fits 24 GB | 21.3 GB · fits 24 GB |
| Q5_K_M | 11.5 GB · fits 16 GB | 12.2 GB · fits 16 GB | 16.2 GB · fits 24 GB |
| Q4_K_M | 9.6 GB · fits 12 GB | 10.3 GB · fits 12 GB | 14.3 GB · fits 16 GB |
Weights at this quant: 8.1 GB at Q4. Every extra 1,000 tokens of context adds about 0.164 GB of KV cache at FP16. Try other settings in the VRAM calculator.
Which GPU to pick for Ministral 3 14B
The smallest card here that loads it at Q4_K_M with 8k of context is the RTX 3060 (12 GB): 10.3 GB needed, an estimated 31 tokens per second, and a tight fit that leaves little for the conversation. For room to work, meaning 32,768 tokens of context while staying under 85% of memory, step up to the RTX 3090 (24 GB), which needs 14.3 GB for that and generates at about 81 tokens per second.
At Q8_0, which is near-lossless, the card is the RTX 3090 (24 GB), needing 17.3 GB. On Apple silicon, a MacBook Pro M4 Max holds it with 32,768 tokens of context at an estimated 47 tokens per second; prompt processing is slower than on Nvidia, which long documents make obvious. Among unified-memory desktops, the Ryzen AI Max+ 395 holds it at 32,768 tokens in 14.3 GB, at an estimated 22 tokens per second. To check a card that is not on this page, or a different context length, use the can-I-run-it checker, which opens on this model.
Speed by GPU, at Q4 and 8k context
| GPU | Memory | Bandwidth | Fits | Est. tokens/s | Feels like |
|---|---|---|---|---|---|
| RTX 3060 12 GB | 12 GB | 360 GB/s | tight | 31 | comfortable for chat |
| RTX 4060 Ti 16 GB | 16 GB | 288 GB/s | yes | 25 | usable, a little slow |
| RTX 4070 12 GB | 12 GB | 504 GB/s | tight | 43 | comfortable for chat |
| RTX 3090 24 GB | 24 GB | 936 GB/s | yes | 81 | faster than you can read |
| RTX 4090 24 GB | 24 GB | 1,008 GB/s | yes | 87 | faster than you can read |
| RTX 5090 32 GB | 32 GB | 1,792 GB/s | yes | ≤ 154 | faster than you can read |
| RTX 6000 Ada 48 GB | 48 GB | 960 GB/s | yes | 83 | faster than you can read |
| L40S 48 GB | 48 GB | 864 GB/s | yes | 74 | faster than you can read |
| RTX PRO 6000 Blackwell 96 GB | 96 GB | 1,792 GB/s | yes | ≤ 154 | faster than you can read |
| A100 80 GB | 80 GB | 2,039 GB/s | yes | ≤ 176 | faster than you can read |
| H100 SXM 80 GB | 80 GB | 3,350 GB/s | yes | ≤ 289 | faster than you can read |
| RTX 5060 Ti 16 GB | 16 GB | 448 GB/s | yes | 39 | comfortable for chat |
| Radeon RX 7900 XTX 24 GB | 24 GB | 960 GB/s | yes | 83 | faster than you can read |
| DGX Spark 128 GB (unified) | 128 GB (126 for the GPU) | 273 GB/s | yes | 24 | usable, a little slow |
| Ryzen AI Max+ 395 128 GB (unified) | 128 GB (96 for the GPU) | 256 GB/s | yes | 22 | usable, a little slow |
Single-stream decode, from memory bandwidth at 70% efficiency. Prompt processing and batching not included. These are ceilings, not measurements: Ministral 3 14B reads only 8.1 GB of weights per token at Q4, so little that the per-token costs the formula leaves out decide much of the real speed. The 4 figures marked ≤ are where a real runtime falls furthest below the number shown. See the speed estimator for other quantisations and Apple hardware, or every card compared if you are choosing hardware rather than a model.
What Ministral 3 14B is good at, and what it is not
Good at
- A 16 GB card: it fits with a useful context, and Mistral ships it in FP8 for cards that support it.
- Function calling and JSON output, which the card lists as design goals.
- European-language work: French, Spanish, German, Italian, Portuguese and Dutch are named explicitly.
Watch out for
- Every layer keeps a full-length cache, so long context costs several times what a 2026 hybrid model pays for the same conversation.
- The 256k window is real, but filling it costs more memory than the weights do.
Similar models
The models that need about the same memory as Ministral 3 14B at Q4_K_M and 8k, which makes them the real alternatives on whatever card you have. The context column is what each extra 1,000 tokens costs, and it is where models of the same size differ most.
| Model | Size | Needs | Per 1k context | Window | Released |
|---|---|---|---|---|---|
| Ministral 3 14B | 14B | 10.3 GB | 0.164 GB | 256k | 2025-12 |
| Qwen3 14B | 14.8B | 10.8 GB | 0.164 GB | 40k | 2025-04 |
| Phi-4 14B | 14.7B | 11 GB | 0.205 GB | 16k | 2024-12 |
| Gemma 3 12B | 12.2B | 8.7 GB | 0.066 GB | 128k | 2025-03 |
| Gemma 4 12B | 12B | 8.2 GB | 0.016 GB | 256k | 2026-05 |
Notes
- The parameter count includes the vision encoder; a text-only GGUF is slightly smaller.
- Native context window: 262,144 tokens. Licence: Apache 2.0, a permissive licence.
- Weights published in December 2025.
- Architecture values from config.json in mistralai/Ministral-3-14B-Instruct-2512 on Hugging Face. Every figure on this page is computed from those values; check them before buying hardware for this model.
- Estimates assume a single conversation on a card that is otherwise free. A desktop on the same GPU takes 0.5 to 2 GB.
Questions
Can I run Ministral 3 14B on a 24 GB card?
Yes, at Q4_K_M and 8k context it needs about 10.3 GB, which fits a 24 GB card with room.
How much VRAM does Ministral 3 14B need at Q8?
About 17.3 GB at 8k context, or 21.3 GB at 32k. Q8 is near-lossless; use it when it fits.
What is the context length of Ministral 3 14B?
262,144 tokens natively. Holding all of it at Q4_K_M takes about 51.9 GB, of which 42.9 GB is context.
How fast is Ministral 3 14B on an RTX 4090?
At most about 87 tokens per second at Q4, single stream, which is faster than you can read. That is a ceiling worked out from the card's memory bandwidth, not a measurement.
See every model compared and which models fit on 8 to 96 GB, or run it without buying the card: Nodegrove attaches a 24, 48 or 96 GB GPU to a workspace that stays saved.