Qwen3.5 4B has 4.7B parameters. It has 32 layers, of which only 8 are full attention (4 KV heads of dimension 256) and keep a cache that grows with the conversation; the other 24 are Gated DeltaNet linear-attention layers with a small fixed state. At the everyday setting, Q4_K_M and 8k context, it needs an estimated 3.7 GB, so the smallest card that fits is 8 GB.
What Qwen3.5 4B is
The small end of Alibaba's Qwen3.5 family from February 2026. It reads images and video as well as text, covers 201 languages, and uses the hybrid architecture that makes long context nearly free in memory terms.
Text, images and video 262,144-token window Apache 2.0 February 2026
Memory by quantisation and context
| Quant | 4,096 ctx | 8,192 ctx | 32,768 ctx |
|---|---|---|---|
| Q8_0 | 5.9 GB · fits 8 GB | 6 GB · fits 8 GB | 6.8 GB · fits 8 GB |
| Q5_K_M | 4.2 GB · fits 8 GB | 4.3 GB · fits 8 GB | 5.1 GB · fits 8 GB |
| Q4_K_M | 3.5 GB · fits 8 GB | 3.7 GB · fits 8 GB | 4.5 GB · fits 8 GB |
Weights at this quant: 2.7 GB at Q4. Every extra 1,000 tokens of context adds about 0.033 GB of KV cache at FP16, on top of a fixed 0.05 GB of recurrent state that does not grow. Try other settings in the VRAM calculator.
Which GPU to pick for Qwen3.5 4B
The smallest card here that loads it at Q4_K_M with 8k of context is the RTX 3060 (12 GB): 3.7 GB needed, an estimated 92 tokens per second. That is only the smallest card with its own page here: the model itself fits any 8 GB card. The same card still has room at 32,768 tokens of context (4.5 GB), so there is no reason to buy above it for this model.
The same card also holds it at Q8_0, in 6 GB. Q8 is near-lossless, so take the higher precision when it costs nothing. On Apple silicon, a MacBook Pro M4 Max holds it with 32,768 tokens of context at an estimated 140 tokens per second; prompt processing is slower than on Nvidia, which long documents make obvious. Among unified-memory desktops, the Ryzen AI Max+ 395 holds it at 32,768 tokens in 4.5 GB, at an estimated 66 tokens per second. To check a card that is not on this page, or a different context length, use the can-I-run-it checker, which opens on this model.
Speed by GPU, at Q4 and 8k context
| GPU | Memory | Bandwidth | Fits | Est. tokens/s | Feels like |
|---|---|---|---|---|---|
| RTX 3060 12 GB | 12 GB | 360 GB/s | yes | 92 | faster than you can read |
| RTX 4060 Ti 16 GB | 16 GB | 288 GB/s | yes | 74 | faster than you can read |
| RTX 4070 12 GB | 12 GB | 504 GB/s | yes | ≤ 129 | faster than you can read |
| RTX 3090 24 GB | 24 GB | 936 GB/s | yes | ≤ 240 | faster than you can read |
| RTX 4090 24 GB | 24 GB | 1,008 GB/s | yes | ≤ 259 | faster than you can read |
| RTX 5090 32 GB | 32 GB | 1,792 GB/s | yes | ≤ 460 | faster than you can read |
| RTX 6000 Ada 48 GB | 48 GB | 960 GB/s | yes | ≤ 247 | faster than you can read |
| L40S 48 GB | 48 GB | 864 GB/s | yes | ≤ 222 | faster than you can read |
| RTX PRO 6000 Blackwell 96 GB | 96 GB | 1,792 GB/s | yes | ≤ 460 | faster than you can read |
| A100 80 GB | 80 GB | 2,039 GB/s | yes | ≤ 524 | faster than you can read |
| H100 SXM 80 GB | 80 GB | 3,350 GB/s | yes | ≤ 860 | faster than you can read |
| RTX 5060 Ti 16 GB | 16 GB | 448 GB/s | yes | ≤ 115 | faster than you can read |
| Radeon RX 7900 XTX 24 GB | 24 GB | 960 GB/s | yes | ≤ 247 | faster than you can read |
| DGX Spark 128 GB (unified) | 128 GB (126 for the GPU) | 273 GB/s | yes | 70 | faster than you can read |
| Ryzen AI Max+ 395 128 GB (unified) | 128 GB (96 for the GPU) | 256 GB/s | yes | 66 | faster than you can read |
Single-stream decode, from memory bandwidth at 70% efficiency. Prompt processing and batching not included. These are ceilings, not measurements: Qwen3.5 4B reads only 2.7 GB of weights per token at Q4, so little that the per-token costs the formula leaves out decide much of the real speed. The 11 figures marked ≤ are where a real runtime falls furthest below the number shown. See the speed estimator for other quantisations and Apple hardware, or every card compared if you are choosing hardware rather than a model.
What Qwen3.5 4B is good at, and what it is not
Good at
- Long documents on a small card: its window is 262,144 tokens and the cache barely grows.
- Describing screenshots, photos and scanned pages locally, on hardware most people already own.
- Anything multilingual at this size.
Watch out for
- It is still a 4B model. The long window lets it read a lot; it does not make it reason like a 27B.
- The hybrid layers need a runtime that implements them. On an old build the memory saving disappears.
Similar models
The models that need about the same memory as Qwen3.5 4B at Q4_K_M and 8k, which makes them the real alternatives on whatever card you have. The context column is what each extra 1,000 tokens costs, and it is where models of the same size differ most.
| Model | Size | Needs | Per 1k context | Window | Released |
|---|---|---|---|---|---|
| Qwen3.5 4B | 4.7B | 3.7 GB | 0.033 GB | 256k | 2026-02 |
| Llama 3.2 3B | 3.2B | 3.4 GB | 0.115 GB | 128k | 2024-09 |
| Llama 3.1 8B | 8B | 6.4 GB | 0.131 GB | 128k | 2024-07 |
| Qwen3 8B | 8.2B | 6.7 GB | 0.147 GB | 40k | 2025-04 |
| Qwen3.5 9B | 9.7B | 6.7 GB | 0.033 GB | 256k | 2026-02 |
Notes
- The parameter count includes the vision encoder; a text-only GGUF is slightly smaller.
- Native context window: 262,144 tokens. Licence: Apache 2.0, a permissive licence.
- Weights published in February 2026.
- Architecture values from config.json in Qwen/Qwen3.5-4B on Hugging Face. Every figure on this page is computed from those values; check them before buying hardware for this model.
- Estimates assume a single conversation on a card that is otherwise free. A desktop on the same GPU takes 0.5 to 2 GB.
Questions
Can I run Qwen3.5 4B on a 24 GB card?
Yes, at Q4_K_M and 8k context it needs about 3.7 GB, which fits a 24 GB card with room.
How much VRAM does Qwen3.5 4B need at Q8?
About 6 GB at 8k context, or 6.8 GB at 32k. Q8 is near-lossless; use it when it fits.
What is the context length of Qwen3.5 4B?
262,144 tokens natively. Holding all of it at Q4_K_M takes about 12 GB, of which 8.6 GB is context.
Why does Qwen3.5 4B need so little memory for long context?
Because most of its layers do not keep a cache that grows. It has 32 layers, of which only 8 are full attention (4 KV heads of dimension 256) and keep a cache that grows with the conversation; the other 24 are Gated DeltaNet linear-attention layers with a small fixed state. At 32k tokens the context costs about 1.1 GB; if all 32 layers kept a conventional full-length cache, the same conversation would cost about 4.3 GB.
How fast is Qwen3.5 4B on an RTX 4090?
At most about 259 tokens per second at Q4, single stream, which is faster than you can read. That is a ceiling worked out from the card's memory bandwidth, not a measurement, and at a figure this high real runtimes land well below it, because the costs the formula leaves out take a growing share of each token.
See every model compared and which models fit on 8 to 96 GB, or run it without buying the card: Nodegrove attaches a 24, 48 or 96 GB GPU to a workspace that stays saved.