Qwen3.8 27B has 27.8B parameters. It has 64 layers, of which only 16 are full attention (4 KV heads of dimension 256) and keep a cache that grows with the conversation; the other 48 are Gated DeltaNet linear-attention layers with a fixed state of about 0.15 GB. At the everyday setting, Q4_K_M and 8k context, it needs an estimated 18 GB, so the smallest card that fits is 24 GB.
What Qwen3.8 27B is
Alibaba's dense 27B from August 2026, and one of the most downloaded open models of the year. A native vision-language model built on the Qwen3.5 hybrid architecture, aimed at coding, research and long multi-step agent work.
Text, images and video 262,144-token window Apache 2.0 August 2026
Reasoning. On by default and can be switched off per request; reasoning depth is tunable with a reasoning-effort setting.
Memory by quantisation and context
| Quant | 4,096 ctx | 8,192 ctx | 32,768 ctx |
|---|---|---|---|
| Q8_0 | 31.6 GB · fits 48 GB | 31.8 GB · fits 48 GB | 33.4 GB · fits 48 GB |
| Q5_K_M | 21.4 GB · fits 24 GB | 21.7 GB · fits 24 GB | 23.3 GB · fits 32 GB |
| Q4_K_M | 17.7 GB · fits 24 GB | 18 GB · fits 24 GB | 19.6 GB · fits 24 GB |
Weights at this quant: 16.1 GB at Q4. Every extra 1,000 tokens of context adds about 0.066 GB of KV cache at FP16, on top of a fixed 0.15 GB of recurrent state that does not grow. Try other settings in the VRAM calculator.
Which GPU to pick for Qwen3.8 27B
The smallest card here that loads it at Q4_K_M with 8k of context is the RTX 3090 (24 GB): 18 GB needed, an estimated 41 tokens per second. The same card still has room at 32,768 tokens of context (19.6 GB), so there is no reason to buy above it for this model.
At Q8_0, which is near-lossless, the card is the L40S (48 GB), needing 31.8 GB. On Apple silicon, a MacBook Pro M4 Max holds it with 32,768 tokens of context at an estimated 24 tokens per second; prompt processing is slower than on Nvidia, which long documents make obvious. Among unified-memory desktops, the Ryzen AI Max+ 395 holds it at 32,768 tokens in 19.6 GB, at an estimated 11 tokens per second. To check a card that is not on this page, or a different context length, use the can-I-run-it checker, which opens on this model.
Speed by GPU, at Q4 and 8k context
| GPU | Memory | Bandwidth | Fits | Est. tokens/s | Feels like |
|---|---|---|---|---|---|
| RTX 3060 12 GB | 12 GB | 360 GB/s | no | – | does not fit |
| RTX 4060 Ti 16 GB | 16 GB | 288 GB/s | no | – | does not fit |
| RTX 4070 12 GB | 12 GB | 504 GB/s | no | – | does not fit |
| RTX 3090 24 GB | 24 GB | 936 GB/s | yes | 41 | comfortable for chat |
| RTX 4090 24 GB | 24 GB | 1,008 GB/s | yes | 44 | comfortable for chat |
| RTX 5090 32 GB | 32 GB | 1,792 GB/s | yes | 78 | faster than you can read |
| RTX 6000 Ada 48 GB | 48 GB | 960 GB/s | yes | 42 | comfortable for chat |
| L40S 48 GB | 48 GB | 864 GB/s | yes | 38 | comfortable for chat |
| RTX PRO 6000 Blackwell 96 GB | 96 GB | 1,792 GB/s | yes | 78 | faster than you can read |
| A100 80 GB | 80 GB | 2,039 GB/s | yes | 89 | faster than you can read |
| H100 SXM 80 GB | 80 GB | 3,350 GB/s | yes | ≤ 145 | faster than you can read |
| RTX 5060 Ti 16 GB | 16 GB | 448 GB/s | no | – | does not fit |
| Radeon RX 7900 XTX 24 GB | 24 GB | 960 GB/s | yes | 42 | comfortable for chat |
| DGX Spark 128 GB (unified) | 128 GB (126 for the GPU) | 273 GB/s | yes | 12 | slow; fine for batch jobs |
| Ryzen AI Max+ 395 128 GB (unified) | 128 GB (96 for the GPU) | 256 GB/s | yes | 11 | slow; fine for batch jobs |
Single-stream decode, from memory bandwidth at 70% efficiency. Prompt processing and batching not included. These are ceilings, not measurements: Qwen3.8 27B reads only 16.1 GB of weights per token at Q4, so little that the per-token costs the formula leaves out decide much of the real speed. The figure marked ≤ is where a real runtime falls furthest below the number shown. See the speed estimator for other quantisations and Apple hardware, or every card compared if you are choosing hardware rather than a model.
What Qwen3.8 27B is good at, and what it is not
Good at
- The 24 GB card. It fits at Q4 with tens of thousands of tokens of context left, where its predecessor, Qwen3 32B, filled the same card with weights.
- Agentic coding and long tasks, which is what the release notes say it was tuned for.
- Reading images, documents and long videos without a second model.
- Apache 2.0.
Watch out for
- Thinking is on by default, and thinking tokens fill the context like any others. Turn it off for quick questions.
- The memory figures assume a runtime with Gated DeltaNet support. On an older build, expect the 2025 arithmetic.
- At Q8 it needs a 48 GB card; there is no 32 GB middle ground.
Similar models
The models that need about the same memory as Qwen3.8 27B at Q4_K_M and 8k, which makes them the real alternatives on whatever card you have. The context column is what each extra 1,000 tokens costs, and it is where models of the same size differ most.
| Model | Size | Needs | Per 1k context | Window | Released |
|---|---|---|---|---|---|
| Qwen3.8 27B | 27.8B | 18 GB | 0.066 GB | 256k | 2026-08 |
| Gemma 3 27B | 27.4B | 18.1 GB | 0.082 GB | 128k | 2025-03 |
| Gemma 4 26B-A4B (MoE) | 25.8B · 3.8B active | 16.4 GB | 0.02 GB | 256k | 2026-03 |
| Nemotron 3.5 Lightning 30B-A3B (MoE) | 31.6B · 3B active | 19.7 GB | 0.006 GB | 256k | 2026-08 |
| Qwen3 30B-A3B (MoE) | 30.5B · 3.3B active | 19.7 GB | 0.098 GB | 40k | 2025-04 |
Notes
- The parameter count includes the vision encoder; a text-only GGUF is slightly smaller.
- Native context window: 262,144 tokens. Licence: Apache 2.0, a permissive licence.
- Weights published in August 2026. It follows Qwen3 32B.
- Architecture values from config.json in Qwen/Qwen3.8-27B on Hugging Face. Every figure on this page is computed from those values; check them before buying hardware for this model.
- Estimates assume a single conversation on a card that is otherwise free. A desktop on the same GPU takes 0.5 to 2 GB.
Questions
Can I run Qwen3.8 27B on a 24 GB card?
Yes, at Q4_K_M and 8k context it needs about 18 GB, which fits a 24 GB card with room.
How much VRAM does Qwen3.8 27B need at Q8?
About 31.8 GB at 8k context, or 33.4 GB at 32k. Q8 is near-lossless; use it when it fits.
What is the context length of Qwen3.8 27B?
262,144 tokens natively. Holding all of it at Q4_K_M takes about 34.6 GB, of which 17.3 GB is context.
Why does Qwen3.8 27B need so little memory for long context?
Because most of its layers do not keep a cache that grows. It has 64 layers, of which only 16 are full attention (4 KV heads of dimension 256) and keep a cache that grows with the conversation; the other 48 are Gated DeltaNet linear-attention layers with a fixed state of about 0.15 GB. At 32k tokens the context costs about 2.3 GB; if all 64 layers kept a conventional full-length cache, the same conversation would cost about 8.6 GB.
Can Qwen3.8 27B run on an RTX 4090 or 3090?
Yes, at Q4_K_M, and with a long context. Only 16 of its 64 layers keep a cache that grows, so the conversation costs a fraction of what it does on a standard 27B to 32B model. The tables above show the exact figures for both cards.
How fast is Qwen3.8 27B on an RTX 4090?
At most about 44 tokens per second at Q4, single stream, which is comfortable for chat. That is a ceiling worked out from the card's memory bandwidth, not a measurement.
See every model compared and which models fit on 8 to 96 GB, or run it without buying the card: Nodegrove attaches a 24, 48 or 96 GB GPU to a workspace that stays saved.