GLM-4.7 Flash 30B-A3B (MoE) has 31.2B parameters with 3B active per token (mixture of experts). It has 47 layers using multi-head latent attention: instead of keys and values for each of its 20 heads, every layer caches one compressed vector of 576 values per token. At the everyday setting, Q4_K_M and 8k context, it needs an estimated 19.8 GB, so the smallest card that fits is 24 GB.
What GLM-4.7 Flash 30B-A3B (MoE) is
Z.ai's 30B-A3B mixture-of-experts model, January 2026, the lightweight member of the GLM-4.7 line. MIT-licensed, text only, with latent attention that keeps its cache small.
Text only 202,752-token window MIT January 2026 3B of 31.2B active
Reasoning. Supports thinking, including a preserved-thinking mode for multi-turn agent work.
Memory by quantisation and context
| Quant | 4,096 ctx | 8,192 ctx | 32,768 ctx |
|---|---|---|---|
| Q8_0 | 35.1 GB · fits 48 GB | 35.3 GB · fits 48 GB | 36.7 GB · fits 48 GB |
| Q5_K_M | 23.8 GB · fits 32 GB | 24 GB · fits 32 GB | 25.3 GB · fits 32 GB |
| Q4_K_M | 19.5 GB · fits 24 GB | 19.8 GB · fits 24 GB | 21.1 GB · fits 24 GB |
Weights at this quant: 18.1 GB at Q4. Every extra 1,000 tokens of context adds about 0.054 GB of KV cache at FP16. Try other settings in the VRAM calculator.
Which GPU to pick for GLM-4.7 Flash 30B-A3B (MoE)
The smallest card here that loads it at Q4_K_M with 8k of context is the RTX 3090 (24 GB): 19.8 GB needed, an estimated 377 tokens per second. For room to work, meaning 32,768 tokens of context while staying under 85% of memory, step up to the RTX 5090 (32 GB), which needs 21.1 GB for that and generates at about 721 tokens per second.
At Q8_0, which is near-lossless, the card is the L40S (48 GB), needing 35.3 GB. On Apple silicon, a MacBook Pro M4 Max holds it with 32,768 tokens of context at an estimated 220 tokens per second; prompt processing is slower than on Nvidia, which long documents make obvious. Among unified-memory desktops, the Ryzen AI Max+ 395 holds it at 32,768 tokens in 21.1 GB, at an estimated 103 tokens per second. To check a card that is not on this page, or a different context length, use the can-I-run-it checker, which opens on this model.
Speed by GPU, at Q4 and 8k context
| GPU | Memory | Bandwidth | Fits | Est. tokens/s | Feels like |
|---|---|---|---|---|---|
| RTX 3060 12 GB | 12 GB | 360 GB/s | no | – | does not fit |
| RTX 4060 Ti 16 GB | 16 GB | 288 GB/s | no | – | does not fit |
| RTX 4070 12 GB | 12 GB | 504 GB/s | no | – | does not fit |
| RTX 3090 24 GB | 24 GB | 936 GB/s | yes | ≤ 377 | faster than you can read |
| RTX 4090 24 GB | 24 GB | 1,008 GB/s | yes | ≤ 406 | faster than you can read |
| RTX 5090 32 GB | 32 GB | 1,792 GB/s | yes | ≤ 721 | faster than you can read |
| RTX 6000 Ada 48 GB | 48 GB | 960 GB/s | yes | ≤ 386 | faster than you can read |
| L40S 48 GB | 48 GB | 864 GB/s | yes | ≤ 348 | faster than you can read |
| RTX PRO 6000 Blackwell 96 GB | 96 GB | 1,792 GB/s | yes | ≤ 721 | faster than you can read |
| A100 80 GB | 80 GB | 2,039 GB/s | yes | ≤ 820 | faster than you can read |
| H100 SXM 80 GB | 80 GB | 3,350 GB/s | yes | ≤ 1348 | faster than you can read |
| RTX 5060 Ti 16 GB | 16 GB | 448 GB/s | no | – | does not fit |
| Radeon RX 7900 XTX 24 GB | 24 GB | 960 GB/s | yes | ≤ 386 | faster than you can read |
| DGX Spark 128 GB (unified) | 128 GB (126 for the GPU) | 273 GB/s | yes | ≤ 110 | faster than you can read |
| Ryzen AI Max+ 395 128 GB (unified) | 128 GB (96 for the GPU) | 256 GB/s | yes | ≤ 103 | faster than you can read |
Single-stream decode, from memory bandwidth at 70% efficiency. Prompt processing and batching not included. These are ceilings, not measurements: GLM-4.7 Flash 30B-A3B (MoE) reads only 1.7 GB of weights per token at Q4, so little that the per-token costs the formula leaves out decide much of the real speed. The 11 figures marked ≤ are where a real runtime falls furthest below the number shown. See the speed estimator for other quantisations and Apple hardware, or every card compared if you are choosing hardware rather than a model.
What GLM-4.7 Flash 30B-A3B (MoE) is good at, and what it is not
Good at
- Agentic and tool-using tasks on a 24 GB card, which is how Z.ai positions it.
- The most permissive licence among the 30B-class mixture-of-experts models here.
Watch out for
- The model card names vLLM and SGLang as its runtimes. Before relying on the context figures here, check that yours implements the latent cache; one that does not stores full keys and values for 20 heads and needs several times more.
- Text only.
Similar models
The models that need about the same memory as GLM-4.7 Flash 30B-A3B (MoE) at Q4_K_M and 8k, which makes them the real alternatives on whatever card you have. The context column is what each extra 1,000 tokens costs, and it is where models of the same size differ most.
| Model | Size | Needs | Per 1k context | Window | Released |
|---|---|---|---|---|---|
| GLM-4.7 Flash 30B-A3B (MoE) | 31.2B · 3B active | 19.8 GB | 0.054 GB | 198k | 2026-01 |
| Qwen3 30B-A3B (MoE) | 30.5B · 3.3B active | 19.7 GB | 0.098 GB | 40k | 2025-04 |
| Nemotron 3.5 Lightning 30B-A3B (MoE) | 31.6B · 3B active | 19.7 GB | 0.006 GB | 256k | 2026-08 |
| Gemma 4 31B | 31.3B | 20.9 GB | 0.082 GB | 256k | 2026-03 |
| Gemma 3 27B | 27.4B | 18.1 GB | 0.082 GB | 128k | 2025-03 |
Notes
- A runtime that does not implement the latent cache stores full keys and values and needs several times more context memory.
- Native context window: 202,752 tokens. Licence: MIT, a permissive licence.
- Weights published in January 2026.
- Architecture values from config.json in zai-org/GLM-4.7-Flash on Hugging Face, and its model card. Every figure on this page is computed from those values; check them before buying hardware for this model.
- Estimates assume a single conversation on a card that is otherwise free. A desktop on the same GPU takes 0.5 to 2 GB.
Questions
Can I run GLM-4.7 Flash 30B-A3B (MoE) on a 24 GB card?
Yes, at Q4_K_M and 8k context it needs about 19.8 GB, which fits a 24 GB card with room.
How much VRAM does GLM-4.7 Flash 30B-A3B (MoE) need at Q8?
About 35.3 GB at 8k context, or 36.7 GB at 32k. Q8 is near-lossless; use it when it fits.
What is the context length of GLM-4.7 Flash 30B-A3B (MoE)?
202,752 tokens natively. Holding all of it at Q4_K_M takes about 30.3 GB, of which 11 GB is context.
Why does GLM-4.7 Flash 30B-A3B (MoE) need so little memory for long context?
Because most of its layers do not keep a cache that grows. It has 47 layers using multi-head latent attention: instead of keys and values for each of its 20 heads, every layer caches one compressed vector of 576 values per token. At 32k tokens the context costs about 1.8 GB; if all 47 layers kept a conventional full-length cache, the same conversation would cost about 31.5 GB.
How fast is GLM-4.7 Flash 30B-A3B (MoE) on an RTX 4090?
At most about 406 tokens per second at Q4, single stream, which is faster than you can read. That is a ceiling worked out from the card's memory bandwidth, not a measurement, and at a figure this high real runtimes land well below it, because the costs the formula leaves out take a growing share of each token.
See every model compared and which models fit on 8 to 96 GB, or run it without buying the card: Nodegrove attaches a 24, 48 or 96 GB GPU to a workspace that stays saved.