Gemma 3 12B has 12.2B parameters, 48 layers and 8 KV heads of dimension 256. At the everyday setting, Q4_K_M and 8k context, it needs an estimated 11.1 GB, so the smallest card that fits is 12 GB.
Memory by quantisation and context
| Quant | 4,096 ctx | 8,192 ctx | 32,768 ctx |
|---|---|---|---|
| Q8_0 | 15.6 GB · fits 24 GB | 17.2 GB · fits 24 GB | 26.8 GB · fits 32 GB |
| Q5_K_M | 11.1 GB · fits 12 GB | 12.7 GB · fits 16 GB | 22.4 GB · fits 24 GB |
| Q4_K_M | 9.5 GB · fits 12 GB | 11.1 GB · fits 12 GB | 20.7 GB · fits 24 GB |
Weights at this quant: 7.1 GB at Q4. Every extra 1,000 tokens of context adds about 0.3932 GB of KV cache at FP16. Try other settings in the VRAM calculator.
Speed by GPU, at Q4 and 8k context
| GPU | Memory | Bandwidth | Fits | Est. tokens/s | Feels like |
|---|---|---|---|---|---|
| RTX 3060 12 GB | 12 GB | 360 GB/s | tight | 36 | comfortable for chat |
| RTX 4060 Ti 16 GB | 16 GB | 288 GB/s | yes | 28 | comfortable for chat |
| RTX 4070 12 GB | 12 GB | 504 GB/s | tight | 50 | comfortable for chat |
| RTX 3090 24 GB | 24 GB | 936 GB/s | yes | 93 | faster than you can read |
| RTX 4090 24 GB | 24 GB | 1,008 GB/s | yes | 100 | faster than you can read |
| RTX 5090 32 GB | 32 GB | 1,792 GB/s | yes | 177 | faster than you can read |
| RTX 6000 Ada 48 GB | 48 GB | 960 GB/s | yes | 95 | faster than you can read |
| L40S 48 GB | 48 GB | 864 GB/s | yes | 85 | faster than you can read |
| RTX PRO 6000 Blackwell 96 GB | 96 GB | 1,792 GB/s | yes | 177 | faster than you can read |
| A100 80 GB | 80 GB | 2,039 GB/s | yes | 202 | faster than you can read |
| H100 SXM 80 GB | 80 GB | 3,352 GB/s | yes | 332 | faster than you can read |
Single-stream decode ceiling from memory bandwidth at 70% efficiency. Prompt processing and batching not included. See the speed estimator for other quantisations and Apple hardware, or every card compared if you are choosing hardware rather than a model.
Notes
- Uses sliding-window attention on most layers; real KV use is lower than this estimate.
- Architecture values from google/gemma-3-12b-it config.json. Verify against the model card before buying hardware for this model.
- Estimates assume a single conversation on a card that is otherwise free. A desktop on the same GPU takes 0.5 to 2 GB.
Questions
Can I run Gemma 3 12B on a 24 GB card?
Yes, at Q4_K_M and 8k context it needs about 11.1 GB, which fits a 24 GB card with room.
How much VRAM does Gemma 3 12B need at Q8?
About 17.2 GB at 8k context, or 26.8 GB at 32k. Q8 is near-lossless; use it when it fits.
How fast is Gemma 3 12B on an RTX 4090?
Roughly 100 tokens per second at Q4, single stream, which is faster than you can read. Real runtimes land within about 20% of this either way.
See how it compares in which models fit on 8 to 96 GB, or run it without buying the card: Nodegrove attaches a 24, 48 or 96 GB GPU to a workspace that stays saved.