Gemma 4 12B: VRAM requirements and which GPUs run it

How much VRAM Gemma 4 12B needs at Q4, Q5 and Q8, the smallest GPU that fits, and expected tokens per second on common cards.

Updated · estimates are labelled as estimates
On this page

Gemma 4 12B has 12B parameters. It has 48 layers: 40 use sliding-window attention and only remember the last 1,024 tokens (8 KV heads of dimension 256), and 8 are global layers with a single KV head of dimension 512. At the everyday setting, Q4_K_M and 8k context, it needs an estimated 8.2 GB, so the smallest card that fits is 12 GB.

What Gemma 4 12B is

Google DeepMind's Gemma 4 12B "Unified", May 2026. Unified means encoder-free: image patches and raw audio are projected straight into the language model instead of passing through separate encoders, so one 12B checkpoint handles text, images, audio and video.

Text, images, audio and video 262,144-token window Apache 2.0 May 2026

Reasoning. Configurable thinking modes, per the model card.

Memory by quantisation and context

Quant4,096 ctx8,192 ctx32,768 ctx
Q8_014.1 GB · fits 16 GB14.2 GB · fits 16 GB14.6 GB · fits 16 GB
Q5_K_M9.8 GB · fits 12 GB9.8 GB · fits 12 GB10.2 GB · fits 12 GB
Q4_K_M8.1 GB · fits 12 GB8.2 GB · fits 12 GB8.6 GB · fits 12 GB

Weights at this quant: 7 GB at Q4. Every extra 1,000 tokens of context adds about 0.016 GB of KV cache at FP16, on top of a fixed 0.336 GB that the sliding-window layers hold once their window is full. Try other settings in the VRAM calculator.

Which GPU to pick for Gemma 4 12B

The smallest card here that loads it at Q4_K_M with 8k of context is the RTX 3060 (12 GB): 8.2 GB needed, an estimated 36 tokens per second. The same card still has room at 32,768 tokens of context (8.6 GB), so there is no reason to buy above it for this model.

At Q8_0, which is near-lossless, the card is the RTX 4060 Ti (16 GB), needing 14.2 GB. On Apple silicon, a MacBook Pro M4 Max holds it with 32,768 tokens of context at an estimated 55 tokens per second; prompt processing is slower than on Nvidia, which long documents make obvious. Among unified-memory desktops, the Ryzen AI Max+ 395 holds it at 32,768 tokens in 8.6 GB, at an estimated 26 tokens per second. To check a card that is not on this page, or a different context length, use the can-I-run-it checker, which opens on this model.

Speed by GPU, at Q4 and 8k context

GPUMemoryBandwidthFitsEst. tokens/sFeels like
RTX 3060 12 GB12 GB360 GB/syes36comfortable for chat
RTX 4060 Ti 16 GB16 GB288 GB/syes29comfortable for chat
RTX 4070 12 GB12 GB504 GB/syes51comfortable for chat
RTX 3090 24 GB24 GB936 GB/syes94faster than you can read
RTX 4090 24 GB24 GB1,008 GB/syes≤ 101faster than you can read
RTX 5090 32 GB32 GB1,792 GB/syes≤ 180faster than you can read
RTX 6000 Ada 48 GB48 GB960 GB/syes97faster than you can read
L40S 48 GB48 GB864 GB/syes87faster than you can read
RTX PRO 6000 Blackwell 96 GB96 GB1,792 GB/syes≤ 180faster than you can read
A100 80 GB80 GB2,039 GB/syes≤ 205faster than you can read
H100 SXM 80 GB80 GB3,350 GB/syes≤ 337faster than you can read
RTX 5060 Ti 16 GB16 GB448 GB/syes45comfortable for chat
Radeon RX 7900 XTX 24 GB24 GB960 GB/syes97faster than you can read
DGX Spark 128 GB (unified)128 GB (126 for the GPU)273 GB/syes27comfortable for chat
Ryzen AI Max+ 395 128 GB (unified)128 GB (96 for the GPU)256 GB/syes26comfortable for chat

Single-stream decode, from memory bandwidth at 70% efficiency. Prompt processing and batching not included. These are ceilings, not measurements: Gemma 4 12B reads only 7 GB of weights per token at Q4, so little that the per-token costs the formula leaves out decide much of the real speed. The 5 figures marked ≤ are where a real runtime falls furthest below the number shown. See the speed estimator for other quantisations and Apple hardware, or every card compared if you are choosing hardware rather than a model.

What Gemma 4 12B is good at, and what it is not

Good at

  • The most modalities per gigabyte of any model here: it is the only one at this size that takes audio.
  • A 12 GB card with real context to spare, because 40 of its 48 layers only remember the last 1,024 tokens.
  • Commercial use without reading custom terms: Gemma 4 moved to Apache 2.0.

Watch out for

  • The sliding-window saving depends on the runtime. One that stores every layer at full length needs far more memory at long context.
  • Audio and video input need a runtime that supports them; many local ones handle text and images first.

Similar models

The models that need about the same memory as Gemma 4 12B at Q4_K_M and 8k, which makes them the real alternatives on whatever card you have. The context column is what each extra 1,000 tokens costs, and it is where models of the same size differ most.

ModelSizeNeedsPer 1k contextWindowReleased
Gemma 4 12B12B8.2 GB0.016 GB256k2026-05
Gemma 3 12B12.2B8.7 GB0.066 GB128k2025-03
Qwen3.5 9B9.7B6.7 GB0.033 GB256k2026-02
Qwen3 8B8.2B6.7 GB0.147 GB40k2025-04
Ministral 3 14B14B10.3 GB0.164 GB256k2025-12

Notes

  • A runtime without a sliding-window cache stores every layer at full length and needs far more than this at long context.
  • Native context window: 262,144 tokens. Licence: Apache 2.0, a permissive licence.
  • Weights published in May 2026. It follows Gemma 3 12B.
  • Architecture values from config.json in google/gemma-4-12B-it on Hugging Face. Every figure on this page is computed from those values; check them before buying hardware for this model.
  • Estimates assume a single conversation on a card that is otherwise free. A desktop on the same GPU takes 0.5 to 2 GB.

Questions

Can I run Gemma 4 12B on a 24 GB card?

Yes, at Q4_K_M and 8k context it needs about 8.2 GB, which fits a 24 GB card with room.

How much VRAM does Gemma 4 12B need at Q8?

About 14.2 GB at 8k context, or 14.6 GB at 32k. Q8 is near-lossless; use it when it fits.

What is the context length of Gemma 4 12B?

262,144 tokens natively. Holding all of it at Q4_K_M takes about 12.4 GB, of which 4.6 GB is context.

Why does Gemma 4 12B need so little memory for long context?

Because most of its layers do not keep a cache that grows. It has 48 layers: 40 use sliding-window attention and only remember the last 1,024 tokens (8 KV heads of dimension 256), and 8 are global layers with a single KV head of dimension 512. At 32k tokens the context costs about 0.9 GB; if all 48 layers kept a conventional full-length cache, the same conversation would cost about 12.9 GB. That saving depends on the runtime implementing the sliding-window cache. llama.cpp and vLLM do; a build that does not will use far more.

Is Gemma 4 12B open source?

The weights are released under Apache 2.0, a permissive licence. That is a change from Gemma 3, which used Google's own Gemma Terms of Use with its own conditions.

How fast is Gemma 4 12B on an RTX 4090?

At most about 101 tokens per second at Q4, single stream, which is faster than you can read. That is a ceiling worked out from the card's memory bandwidth, not a measurement, and at a figure this high real runtimes land well below it, because the costs the formula leaves out take a growing share of each token.

See every model compared and which models fit on 8 to 96 GB, or run it without buying the card: Nodegrove attaches a 24, 48 or 96 GB GPU to a workspace that stays saved.