Qwen3-Coder-Next 80B-A3B (MoE): VRAM requirements and which GPUs run it

How much VRAM Qwen3-Coder-Next 80B-A3B (MoE) needs at Q4, Q5 and Q8, the smallest GPU that fits, and expected tokens per second on common cards.

Updated · estimates are labelled as estimates
On this page

Qwen3-Coder-Next 80B-A3B (MoE) has 79.7B parameters with 3B active per token (mixture of experts). It has 48 layers, of which only 12 are full attention (2 KV heads of dimension 256) and keep a cache that grows with the conversation; the other 36 are Gated DeltaNet linear-attention layers with a small fixed state. At the everyday setting, Q4_K_M and 8k context, it needs an estimated 48.9 GB, so the smallest card that fits is 80 GB.

What Qwen3-Coder-Next 80B-A3B (MoE) is

Alibaba's coding-agent model, early 2026: 80B of weights across 512 experts, with 3B read per token. Built for coding agents and local development, and designed to plug into CLI and IDE tools.

Text only 262,144-token window Apache 2.0 January 2026 3B of 79.7B active

Reasoning. Non-thinking only: it answers directly and never emits reasoning blocks.

Memory by quantisation and context

Quant4,096 ctx8,192 ctx32,768 ctx
Q8_088.5 GB · fits 96 GB88.6 GB · fits 96 GB89.2 GB · fits 96 GB
Q5_K_M59.5 GB · fits 80 GB59.6 GB · fits 80 GB60.2 GB · fits 80 GB
Q4_K_M48.8 GB · fits 80 GB48.9 GB · fits 80 GB49.5 GB · fits 80 GB

Weights at this quant: 46.2 GB at Q4. Every extra 1,000 tokens of context adds about 0.025 GB of KV cache at FP16, on top of a fixed 0.08 GB of recurrent state that does not grow. Try other settings in the VRAM calculator.

Which GPU to pick for Qwen3-Coder-Next 80B-A3B (MoE)

The smallest card here that loads it at Q4_K_M with 8k of context is the A100 (80 GB): 48.9 GB needed, an estimated 820 tokens per second. The same card still has room at 32,768 tokens of context (49.5 GB), so there is no reason to buy above it for this model.

At Q8_0, which is near-lossless, the card is the RTX PRO 6000 Blackwell (96 GB), needing 88.6 GB. On Apple silicon, a MacBook Pro M4 Max holds it with 32,768 tokens of context at an estimated 220 tokens per second; prompt processing is slower than on Nvidia, which long documents make obvious. Among unified-memory desktops, the Ryzen AI Max+ 395 holds it at 32,768 tokens in 49.5 GB, at an estimated 103 tokens per second. To check a card that is not on this page, or a different context length, use the can-I-run-it checker, which opens on this model.

Speed by GPU, at Q4 and 8k context

GPUMemoryBandwidthFitsEst. tokens/sFeels like
RTX 3060 12 GB12 GB360 GB/sno–does not fit
RTX 4060 Ti 16 GB16 GB288 GB/sno–does not fit
RTX 4070 12 GB12 GB504 GB/sno–does not fit
RTX 3090 24 GB24 GB936 GB/sno–does not fit
RTX 4090 24 GB24 GB1,008 GB/sno–does not fit
RTX 5090 32 GB32 GB1,792 GB/sno–does not fit
RTX 6000 Ada 48 GB48 GB960 GB/sno–does not fit
L40S 48 GB48 GB864 GB/sno–does not fit
RTX PRO 6000 Blackwell 96 GB96 GB1,792 GB/syes≤ 721faster than you can read
A100 80 GB80 GB2,039 GB/syes≤ 820faster than you can read
H100 SXM 80 GB80 GB3,350 GB/syes≤ 1348faster than you can read
RTX 5060 Ti 16 GB16 GB448 GB/sno–does not fit
Radeon RX 7900 XTX 24 GB24 GB960 GB/sno–does not fit
DGX Spark 128 GB (unified)128 GB (126 for the GPU)273 GB/syes≤ 110faster than you can read
Ryzen AI Max+ 395 128 GB (unified)128 GB (96 for the GPU)256 GB/syes≤ 103faster than you can read

Single-stream decode, from memory bandwidth at 70% efficiency. Prompt processing and batching not included. These are ceilings, not measurements: Qwen3-Coder-Next 80B-A3B (MoE) reads only 1.7 GB of weights per token at Q4, so little that the per-token costs the formula leaves out decide much of the real speed. The 5 figures marked ≤ are where a real runtime falls furthest below the number shown. See the speed estimator for other quantisations and Apple hardware, or every card compared if you are choosing hardware rather than a model.

What Qwen3-Coder-Next 80B-A3B (MoE) is good at, and what it is not

Good at

  • A coding agent that answers at small-model speed, on an 80 or 96 GB card or a large-memory Mac.
  • Repository-scale context: a 262,144-token window with 12 cache-growing layers out of 48.
  • Agent loops, where no thinking preamble means no wasted tokens per step.

Watch out for

  • It is the memory that costs, not the compute. It does not fit a 48 GB card at Q4.
  • A specialist. It is not the model for general conversation.
  • The hybrid layers need a current runtime.

Similar models

The models that need about the same memory as Qwen3-Coder-Next 80B-A3B (MoE) at Q4_K_M and 8k, which makes them the real alternatives on whatever card you have. The context column is what each extra 1,000 tokens costs, and it is where models of the same size differ most.

ModelSizeNeedsPer 1k contextWindowReleased
Qwen3-Coder-Next 80B-A3B (MoE)79.7B · 3B active48.9 GB0.025 GB256k2026-01
DeepSeek-R1 Distill Llama 70B70.6B45.8 GB0.328 GB128k2025-01
Llama 3.3 70B70.6B45.8 GB0.328 GB128k2024-12
gpt-oss 120B117B · 5.1B active71.4 GB0.037 GB128k2025-08
Mistral Small 4 119B (MoE)119B · 6.5B active72.5 GB0.023 GB256k2026-03

Notes

  • All 512 experts stay in memory; 3B are active per token, so it answers like a small model once it is loaded.
  • Native context window: 262,144 tokens. Licence: Apache 2.0, a permissive licence.
  • This is a coding model, tuned for writing and editing code and for agentic tools that read whole repositories, which is where a long context earns its memory.
  • Weights published in January 2026. It follows Qwen2.5 Coder 32B.
  • Architecture values from config.json in Qwen/Qwen3-Coder-Next on Hugging Face, and its model card. Every figure on this page is computed from those values; check them before buying hardware for this model.
  • Estimates assume a single conversation on a card that is otherwise free. A desktop on the same GPU takes 0.5 to 2 GB.

Questions

Can I run Qwen3-Coder-Next 80B-A3B (MoE) on a 24 GB card?

Not comfortably. At Q4_K_M and 8k context it needs about 48.9 GB. The smallest tier that fits is 80 GB.

How much VRAM does Qwen3-Coder-Next 80B-A3B (MoE) need at Q8?

About 88.6 GB at 8k context, or 89.2 GB at 32k. Q8 is near-lossless; use it when it fits.

What is the context length of Qwen3-Coder-Next 80B-A3B (MoE)?

262,144 tokens natively. Holding all of it at Q4_K_M takes about 55.1 GB, of which 6.5 GB is context.

Why does Qwen3-Coder-Next 80B-A3B (MoE) need so little memory for long context?

Because most of its layers do not keep a cache that grows. It has 48 layers, of which only 12 are full attention (2 KV heads of dimension 256) and keep a cache that grows with the conversation; the other 36 are Gated DeltaNet linear-attention layers with a small fixed state. At 32k tokens the context costs about 0.9 GB; if all 48 layers kept a conventional full-length cache, the same conversation would cost about 3.2 GB.

How fast is Qwen3-Coder-Next 80B-A3B (MoE) on an RTX 4090?

It does not fit a 4090 at Q4 and 8k context, so speed would collapse to CPU offloading. Use a larger card or a smaller quantisation.

See every model compared and which models fit on 8 to 96 GB, or run it without buying the card: Nodegrove attaches a 24, 48 or 96 GB GPU to a workspace that stays saved.