Qwen3 8B has 8.2B parameters, 36 layers and 8 KV heads of dimension 128. At the everyday setting, Q4_K_M and 8k context, it needs an estimated 6.7 GB, so the smallest card that fits is 8 GB.
Memory by quantisation and context
| Quant | 4,096 ctx | 8,192 ctx | 32,768 ctx |
|---|---|---|---|
| Q8_0 | 10.1 GB · fits 12 GB | 10.7 GB · fits 12 GB | 14.4 GB · fits 16 GB |
| Q5_K_M | 7.2 GB · fits 8 GB | 7.8 GB · fits 12 GB | 11.4 GB · fits 12 GB |
| Q4_K_M | 6.1 GB · fits 8 GB | 6.7 GB · fits 8 GB | 10.3 GB · fits 12 GB |
Weights at this quant: 4.8 GB at Q4. Every extra 1,000 tokens of context adds about 0.1475 GB of KV cache at FP16. Try other settings in the VRAM calculator.
Speed by GPU, at Q4 and 8k context
| GPU | Memory | Bandwidth | Fits | Est. tokens/s | Feels like |
|---|---|---|---|---|---|
| RTX 3060 12 GB | 12 GB | 360 GB/s | yes | 53 | comfortable for chat |
| RTX 4060 Ti 16 GB | 16 GB | 288 GB/s | yes | 42 | comfortable for chat |
| RTX 4070 12 GB | 12 GB | 504 GB/s | yes | 74 | faster than you can read |
| RTX 3090 24 GB | 24 GB | 936 GB/s | yes | 138 | faster than you can read |
| RTX 4090 24 GB | 24 GB | 1,008 GB/s | yes | 148 | faster than you can read |
| RTX 5090 32 GB | 32 GB | 1,792 GB/s | yes | 264 | faster than you can read |
| RTX 6000 Ada 48 GB | 48 GB | 960 GB/s | yes | 141 | faster than you can read |
| L40S 48 GB | 48 GB | 864 GB/s | yes | 127 | faster than you can read |
| RTX PRO 6000 Blackwell 96 GB | 96 GB | 1,792 GB/s | yes | 264 | faster than you can read |
| A100 80 GB | 80 GB | 2,039 GB/s | yes | 300 | faster than you can read |
| H100 SXM 80 GB | 80 GB | 3,352 GB/s | yes | 493 | faster than you can read |
Single-stream decode ceiling from memory bandwidth at 70% efficiency. Prompt processing and batching not included. See the speed estimator for other quantisations and Apple hardware, or every card compared if you are choosing hardware rather than a model.
Notes
- Architecture values from Qwen/Qwen3-8B config.json. Verify against the model card before buying hardware for this model.
- Estimates assume a single conversation on a card that is otherwise free. A desktop on the same GPU takes 0.5 to 2 GB.
Questions
Can I run Qwen3 8B on a 24 GB card?
Yes, at Q4_K_M and 8k context it needs about 6.7 GB, which fits a 24 GB card with room.
How much VRAM does Qwen3 8B need at Q8?
About 10.7 GB at 8k context, or 14.4 GB at 32k. Q8 is near-lossless; use it when it fits.
How fast is Qwen3 8B on an RTX 4090?
Roughly 148 tokens per second at Q4, single stream, which is faster than you can read. Real runtimes land within about 20% of this either way.
See how it compares in which models fit on 8 to 96 GB, or run it without buying the card: Nodegrove attaches a 24, 48 or 96 GB GPU to a workspace that stays saved.