Devstral Small 2 24B: VRAM requirements and which GPUs run it

How much VRAM Devstral Small 2 24B needs at Q4, Q5 and Q8, the smallest GPU that fits, and expected tokens per second on common cards.

Updated · estimates are labelled as estimates
On this page

Devstral Small 2 24B has 24B parameters, 40 layers and 8 KV heads of dimension 128. At the everyday setting, Q4_K_M and 8k context, it needs an estimated 16.3 GB, so the smallest card that fits is 24 GB.

What Devstral Small 2 24B is

Mistral's agentic coding model, December 2025. It is built to explore a codebase with tools, edit several files and drive a software-engineering agent, and Mistral released its Vibe CLI alongside it.

Text and images 262,144-token window Apache 2.0 December 2025

Reasoning. Answers directly.

Memory by quantisation and context

Quant4,096 ctx8,192 ctx32,768 ctx
Q8_027.6 GB · fits 32 GB28.3 GB · fits 32 GB32.3 GB · fits 48 GB
Q5_K_M18.9 GB · fits 24 GB19.6 GB · fits 24 GB23.6 GB · fits 32 GB
Q4_K_M15.6 GB · fits 24 GB16.3 GB · fits 24 GB20.3 GB · fits 24 GB

Weights at this quant: 13.9 GB at Q4. Every extra 1,000 tokens of context adds about 0.164 GB of KV cache at FP16. Try other settings in the VRAM calculator.

Which GPU to pick for Devstral Small 2 24B

The smallest card here that loads it at Q4_K_M with 8k of context is the RTX 3090 (24 GB): 16.3 GB needed, an estimated 47 tokens per second. The same card still has room at 32,768 tokens of context (20.3 GB), so there is no reason to buy above it for this model.

At Q8_0, which is near-lossless, the card is the RTX 5090 (32 GB), needing 28.3 GB. On Apple silicon, a MacBook Pro M4 Max holds it with 32,768 tokens of context at an estimated 27 tokens per second; prompt processing is slower than on Nvidia, which long documents make obvious. Among unified-memory desktops, the Ryzen AI Max+ 395 holds it at 32,768 tokens in 20.3 GB, at an estimated 13 tokens per second. To check a card that is not on this page, or a different context length, use the can-I-run-it checker, which opens on this model.

Speed by GPU, at Q4 and 8k context

GPUMemoryBandwidthFitsEst. tokens/sFeels like
RTX 3060 12 GB12 GB360 GB/sno–does not fit
RTX 4060 Ti 16 GB16 GB288 GB/sno–does not fit
RTX 4070 12 GB12 GB504 GB/sno–does not fit
RTX 3090 24 GB24 GB936 GB/syes47comfortable for chat
RTX 4090 24 GB24 GB1,008 GB/syes51comfortable for chat
RTX 5090 32 GB32 GB1,792 GB/syes90faster than you can read
RTX 6000 Ada 48 GB48 GB960 GB/syes48comfortable for chat
L40S 48 GB48 GB864 GB/syes43comfortable for chat
RTX PRO 6000 Blackwell 96 GB96 GB1,792 GB/syes90faster than you can read
A100 80 GB80 GB2,039 GB/syes≤ 103faster than you can read
H100 SXM 80 GB80 GB3,350 GB/syes≤ 168faster than you can read
RTX 5060 Ti 16 GB16 GB448 GB/sno–does not fit
Radeon RX 7900 XTX 24 GB24 GB960 GB/syes48comfortable for chat
DGX Spark 128 GB (unified)128 GB (126 for the GPU)273 GB/syes14usable, a little slow
Ryzen AI Max+ 395 128 GB (unified)128 GB (96 for the GPU)256 GB/syes13usable, a little slow

Single-stream decode, from memory bandwidth at 70% efficiency. Prompt processing and batching not included. These are ceilings, not measurements: Devstral Small 2 24B reads only 13.9 GB of weights per token at Q4, so little that the per-token costs the formula leaves out decide much of the real speed. The 2 figures marked ≤ are where a real runtime falls furthest below the number shown. See the speed estimator for other quantisations and Apple hardware, or every card compared if you are choosing hardware rather than a model.

What Devstral Small 2 24B is good at, and what it is not

Good at

  • A coding agent on a single 24 GB card, which is the use Mistral names on the card.
  • Reading screenshots and diagrams as part of a coding task.
  • Apache 2.0, for a coding model you can ship inside a product.

Watch out for

  • It is a specialist. For general conversation a general model of the same size is the better choice.
  • Every layer keeps a full-length cache, so long context costs several times what a 2026 hybrid model pays for the same conversation. Repository-scale context is exactly where that bites.

Similar models

The models that need about the same memory as Devstral Small 2 24B at Q4_K_M and 8k, which makes them the real alternatives on whatever card you have. The context column is what each extra 1,000 tokens costs, and it is where models of the same size differ most.

ModelSizeNeedsPer 1k contextWindowReleased
Devstral Small 2 24B24B16.3 GB0.164 GB256k2025-12
Mistral Small 3.1 24B24B16.3 GB0.164 GB128k2025-03
Gemma 4 26B-A4B (MoE)25.8B · 3.8B active16.4 GB0.02 GB256k2026-03
Qwen3.8 27B27.8B18 GB0.066 GB256k2026-08
Gemma 3 27B27.4B18.1 GB0.082 GB128k2025-03

Notes

  • Mistral's coding model. The parameter count includes the vision encoder; a text-only GGUF is slightly smaller.
  • Native context window: 262,144 tokens. Licence: Apache 2.0, a permissive licence.
  • This is a coding model, tuned for writing and editing code and for agentic tools that read whole repositories, which is where a long context earns its memory.
  • Weights published in December 2025.
  • Architecture values from config.json in mistralai/Devstral-Small-2-24B-Instruct-2512 on Hugging Face. Every figure on this page is computed from those values; check them before buying hardware for this model.
  • Estimates assume a single conversation on a card that is otherwise free. A desktop on the same GPU takes 0.5 to 2 GB.

Questions

Can I run Devstral Small 2 24B on a 24 GB card?

Yes, at Q4_K_M and 8k context it needs about 16.3 GB, which fits a 24 GB card with room.

How much VRAM does Devstral Small 2 24B need at Q8?

About 28.3 GB at 8k context, or 32.3 GB at 32k. Q8 is near-lossless; use it when it fits.

What is the context length of Devstral Small 2 24B?

262,144 tokens natively. Holding all of it at Q4_K_M takes about 57.9 GB, of which 42.9 GB is context.

How fast is Devstral Small 2 24B on an RTX 4090?

At most about 51 tokens per second at Q4, single stream, which is comfortable for chat. That is a ceiling worked out from the card's memory bandwidth, not a measurement.

See every model compared and which models fit on 8 to 96 GB, or run it without buying the card: Nodegrove attaches a 24, 48 or 96 GB GPU to a workspace that stays saved.