What LLMs can a DGX Spark run?

Which open models run on a DGX Spark: what fits at Q4 and Q8, estimated tokens per second, how much context fits, and where the card stops.

Updated · estimates are labelled as estimates
On this page

128 GB unified 273 GB/s 140 W 2025 Sold new

DGX Spark (128 GB) addresses 126 GB of its 128 GB of unified memory, worked out from what NVIDIA documents (see the sources below), at 273 GB/s. Of the 29 open models tracked on this site, it runs 29 at Q4_K_M with 8k of context. The largest is Mistral Small 4 119B (MoE), needing about 72.5 GB and generating an estimated 51 tokens per second — comfortable for chat.

NVIDIA's small Linux box with 128 GB the GPU can share: large models fit, but they are read at about a quarter of an RTX 4090's speed.

What a DGX Spark runs, model by model

Every model at Q4_K_M with 8k of context, the setting most people actually use. This card has 119.7 GB to spend once the 5% safety margin comes off its 126 GB, and it reads that memory at 273 GB/s. Those two numbers decide the last two columns: the first says what loads, the second says how fast it answers. Both are estimates from stated formulas, and the VRAM calculator shows the working.

ModelSizeNeedsEst. tokens/sOn this card
Llama 3.2 3B 3.2B 3.4 GB ≤ 103 Fits
Qwen3.5 4B 4.7B 3.7 GB 70 Fits
Llama 3.1 8B 8B 6.4 GB 41 Fits
Qwen3 8B 8.2B 6.7 GB 40 Fits
Qwen3.5 9B 9.7B 6.7 GB 34 Fits
Gemma 4 12B 12B 8.2 GB 27 Fits
Gemma 3 12B 12.2B 8.7 GB 27 Fits
Ministral 3 14B 14B 10.3 GB 24 Fits
Qwen3 14B 14.8B 10.8 GB 22 Fits
Phi-4 14B 14.7B 11 GB 22 Fits
gpt-oss 20B 21B · 3.6B active 13.4 GB 92 Fits
Devstral Small 2 24B 24B 16.3 GB 14 Fits
Mistral Small 3.1 24B 24B 16.3 GB 14 Fits
Gemma 4 26B-A4B (MoE) 25.8B · 3.8B active 16.4 GB 87 Fits
Gemma 3 27B 27.4B 18.1 GB 12 Fits
Qwen3.8 27B 27.8B 18 GB 12 Fits
Qwen3 30B-A3B (MoE) 30.5B · 3.3B active 19.7 GB ≤ 100 Fits
GLM-4.7 Flash 30B-A3B (MoE) 31.2B · 3B active 19.8 GB ≤ 110 Fits
Gemma 4 31B 31.3B 20.9 GB 11 Fits
Nemotron 3.5 Lightning 30B-A3B (MoE) 31.6B · 3B active 19.7 GB ≤ 110 Fits
Qwen3 32B 32.8B 22.4 GB 10 Fits
Qwen2.5 Coder 32B 32.8B 22.4 GB 10 Fits
DeepSeek-R1 Distill Qwen 32B 32.8B 22.4 GB 10 Fits
Qwen3.6 35B-A3B (MoE) 36B · 3B active 22.4 GB ≤ 110 Fits
Llama 3.3 70B 70.6B 45.8 GB 5 Fits
DeepSeek-R1 Distill Llama 70B 70.6B 45.8 GB 5 Fits
Qwen3-Coder-Next 80B-A3B (MoE) 79.7B · 3B active 48.9 GB ≤ 110 Fits
gpt-oss 120B 117B · 5.1B active 71.4 GB 65 Fits
Mistral Small 4 119B (MoE) 119B · 6.5B active 72.5 GB 51 Fits

On this card that means anything at or under 119.7 GB counts as fitting, and anything above 107.1 GB is marked tight: it loads, but there is little left for the conversation, and a desktop on the same 128 GB can push it over.

The model to actually run on it

Mistral Small 4 119B (MoE) is the best use of this card: 72.5 GB of the 126 GB available, an estimated 51 tokens per second, comfortable for chat, and 131,072 tokens of context still available. It is a mixture-of-experts model, so all 119B parameters sit in memory but only 6.5B are read per token — which is why it is quick for its size.

How this is chosen, since no benchmark is quoted: models are compared by size, a mixture-of-experts model counting at the geometric mean of its total and active parameters (a rule of thumb, not a measurement). The pick is the biggest class that fits with memory to spare and answers at 25 tokens per second or better, and within that class the most recent general-purpose release; coding and reasoning specialists are listed in the table but not recommended by default.

If quality matters more than parameter count, Llama 3.3 70B fits at Q8_0 in about 81 GB. Q8 is near-lossless; a smaller model at Q8 often beats a larger one squeezed into Q3.

How much context actually fits

Model size is the question people ask; context is the one that bites. The KV cache grows linearly with the conversation, so a model that loads comfortably can still run out of memory halfway through a long document. On 126 GB, the context you get is whatever the weights leave behind — which is why two cards that run the same model can feel completely different in use. These are the largest contexts this card holds, at Q4_K_M weights.

ModelMax context, FP16 KVWith Q8 KV cache
Mistral Small 4 119B (MoE) 131,072 tokens 131,072 tokens
Llama 3.3 70B 131,072 tokens 131,072 tokens
Qwen3 32B 40,960 tokens · model limit 40,960 tokens · model limit
Gemma 4 31B 131,072 tokens 131,072 tokens

Every model in this table already reaches its limit at FP16, so quantising the KV cache buys nothing here. Capped at 128k tokens, or at the model's own native window where that is smaller, marked "model limit": free memory beyond that point buys nothing. Rows that hold far more context than their size suggests are hybrid, sliding-window or latent-attention models, which cache only a fraction of what a standard transformer does; each model page shows the working.

Where this card stops

Nothing in this list is out of reach. Every model tracked here fits at Q4_K_M with 8k of context, so the ceiling you will meet is context length and batch size, not parameter count.

The honest take

It holds models no consumer card can, and it runs CUDA, which a Mac and an AMD box do not. The catch is 273 GB/s. A dense 70B model fits comfortably and crawls; mixture-of-experts models, which keep many parameters but read only a few billion per token, are what this machine is for. Think of it as a CUDA development box with a lot of memory, not a fast one.

What it is good at

  • CUDA on 128 GB of unified memory: the NVIDIA software stack, from vLLM to TensorRT-LLM, runs as it does on a datacenter GPU.
  • Blackwell, so FP4 weight formats run natively.
  • A 150 mm box on a 240 W power supply, quiet enough for a desk.

What it is not

  • 273 GB/s: dense models above 30B generate slowly however well they fit.
  • An Arm Linux machine running DGX OS. Tools need ARM64 builds; Windows software does not run.
  • nvidia-smi reports memory use as "Not Supported" and cudaMemGetInfo under-reports, which makes the memory look missing when it is not.

The thing people get wrong: NVIDIA's "up to 200 billion parameters" assumes 4-bit weights, as the datasheet's own footnote says. At Q8 the same memory holds roughly half that.

Buy or rent

Buy it for CUDA and capacity in a small box, not for speed: a used 3090 generates faster on anything that fits in 24 GB. If the large models are an occasional experiment, renting a large card for those hours costs less than a machine that runs them slowly every day.

Retail prices drift, and a figure written today would be wrong within a month, so this page does not carry one. Put the figure you are actually looking at, with your electricity rate and your real monthly hours, into the build vs rent calculator — it already knows this card draws 140 W under load. It returns the month owning becomes cheaper, or tells you it never does. Hours per month decides it far more often than the price does.

Nodegrove is the rented side of that comparison: a workspace that stays saved, with a GPU attached only while you are using it. It is the right answer when your hours are low or your needs change; buying is the right answer when they are high and stable.

Compared with the alternatives

  • Ryzen AI Max+ 395 — 96 GB · 256 GB/s · best fit Mistral Small 4 119B (MoE) at ~48 tok/s. AMD's 128 GB laptop and mini-PC chip: up to 96 GB for the GPU on Windows, and an x86 route to large mixture-of-experts models.
  • MacBook Pro M5 Max — 96 GB · 614 GB/s · best fit Qwen3.8 27B at ~27 tok/s. The laptop for large models in 2026: 128 GB at 614 GB/s, if you buy the 40-core GPU.
  • RTX 5090 — 32 GB · 1,792 GB/s · best fit Qwen3.8 27B at ~78 tok/s. The first consumer card whose memory bandwidth changes what is comfortable.

The unified-memory machines are compared head to head, on the same models, in DGX Spark vs Strix Halo vs Mac Studio.

All 20 cards are compared side by side on the GPU index, and the same numbers from the model's point of view are on each model fit table. To check one specific pairing rather than read a table, use the can-I-run-it checker.

Questions

Can a DGX Spark run a 70B model?

Yes. Llama 3.3 70B at Q4_K_M with 8k of context needs about 45.8 GB, inside the 126 GB this card can address, and it generates an estimated 5 tokens per second — too slow for interactive use.

What is the best model to run on a DGX Spark?

Mistral Small 4 119B (MoE). At Q4_K_M and 8k context it needs about 72.5 GB of the 126 GB available and generates an estimated 51 tokens per second, which is comfortable for chat. If quality matters more than size, Llama 3.3 70B fits at Q8 in about 81 GB.

How many tokens per second does a DGX Spark generate?

It depends entirely on the size of the model, because generating a token means reading every active weight out of memory. At 273 GB/s and 70% efficiency, this card produces an estimated 41 tokens per second on an 8B model at Q4, and about 51 on the largest model it holds, Mistral Small 4 119B (MoE). Reading speed is roughly 5 to 8 words per second, so anything above 25 feels immediate. Treat 41 as a ceiling on 273 GB/s rather than a measurement; the higher a figure, the further real runtimes fall below it.

Is 126 GB enough for running LLMs locally?

It runs 29 of the 29 models tracked here at Q4_K_M with 8k of context, up to 119B parameters. The honest test is not the model list but the context: Mistral Small 4 119B (MoE) on this card holds about 131,072 tokens before memory runs out.

DGX Spark or Ryzen AI Max+ 395 for local models?

Both address about the same memory, so they run the same models. On speed, this card is faster: 273 GB/s against 256 GB/s, and bandwidth is what sets chat speed.

How much of the DGX Spark's 128 GB can the GPU use?

NVIDIA does not publish a figure. The GPU shares the whole pool with the Arm CPU and DGX OS, less a 2 GB display reserve (4 GB is an option in the BIOS), so what a model can use is whatever the system leaves free. This site plans around 126 GB and keeps a 5% margin for the operating system. Do not judge it by nvidia-smi, which shows memory use as "Not Supported" on this machine.

DGX Spark or Mac Studio for local LLMs?

The Spark runs CUDA, so the NVIDIA-only half of the ecosystem, including most fine-tuning and many image and video tools, works on it. A Mac Studio M5 Ultra reads memory at 1,200 GB/s against 273 GB/s and goes up to 512 GB, so it generates faster and holds more, without CUDA. For chatting with big models, the Mac; for building on NVIDIA hardware or fine-tuning, the Spark.

  • Memory: weights (parameters × bytes per parameter) + KV cache (2 × layers × KV heads × head dim × context × bytes for a standard transformer; sliding-window, hybrid and latent-attention models counted as they actually cache) + 0.5 GB runtime + 4% of weights. Full derivation in the VRAM calculator.
  • Speed: 70% × bandwidth ÷ active weight bytes, single stream, no batching, prompt processing excluded. A ceiling, not a measurement: the 6 of 29 figures marked ≤ are where real runtimes fall furthest below the number, because so few weights are read per token. See the tokens-per-second estimator.
  • Card specification: 128 GB, 273 GB/s, 140 W — manufacturer figures; see NVIDIA's own page for this machine. Unified memory: NVIDIA publishes no GPU share: the GPU can use whatever of the 128 GB the operating system leaves free, less a 2 GB display reserve. Figures here use 126 GB; the operating system comes out of the 5% margin every fit keeps, so run the largest models without a desktop session. Source: NVIDIA's DGX Spark release notes (the 2 GB display reserve).
  • Model architecture values come from each model's published config; the self-hosted LLM guide explains what each one changes.