128 GB unified 273 GB/s 140 W 2025 Sold new
DGX Spark (128 GB) addresses 126 GB of its 128 GB of unified memory, worked out from what NVIDIA documents (see the sources below), at 273 GB/s. Of the 29 open models tracked on this site, it runs 29 at Q4_K_M with 8k of context. The largest is Mistral Small 4 119B (MoE), needing about 72.5 GB and generating an estimated 51 tokens per second — comfortable for chat.
NVIDIA's small Linux box with 128 GB the GPU can share: large models fit, but they are read at about a quarter of an RTX 4090's speed.
What a DGX Spark runs, model by model
Every model at Q4_K_M with 8k of context, the setting most people actually use. This card has 119.7 GB to spend once the 5% safety margin comes off its 126 GB, and it reads that memory at 273 GB/s. Those two numbers decide the last two columns: the first says what loads, the second says how fast it answers. Both are estimates from stated formulas, and the VRAM calculator shows the working.
| Model | Size | Needs | Est. tokens/s | On this card |
|---|---|---|---|---|
| Llama 3.2 3B | 3.2B | 3.4 GB | ≤ 103 | Fits |
| Qwen3.5 4B | 4.7B | 3.7 GB | 70 | Fits |
| Llama 3.1 8B | 8B | 6.4 GB | 41 | Fits |
| Qwen3 8B | 8.2B | 6.7 GB | 40 | Fits |
| Qwen3.5 9B | 9.7B | 6.7 GB | 34 | Fits |
| Gemma 4 12B | 12B | 8.2 GB | 27 | Fits |
| Gemma 3 12B | 12.2B | 8.7 GB | 27 | Fits |
| Ministral 3 14B | 14B | 10.3 GB | 24 | Fits |
| Qwen3 14B | 14.8B | 10.8 GB | 22 | Fits |
| Phi-4 14B | 14.7B | 11 GB | 22 | Fits |
| gpt-oss 20B | 21B · 3.6B active | 13.4 GB | 92 | Fits |
| Devstral Small 2 24B | 24B | 16.3 GB | 14 | Fits |
| Mistral Small 3.1 24B | 24B | 16.3 GB | 14 | Fits |
| Gemma 4 26B-A4B (MoE) | 25.8B · 3.8B active | 16.4 GB | 87 | Fits |
| Gemma 3 27B | 27.4B | 18.1 GB | 12 | Fits |
| Qwen3.8 27B | 27.8B | 18 GB | 12 | Fits |
| Qwen3 30B-A3B (MoE) | 30.5B · 3.3B active | 19.7 GB | ≤ 100 | Fits |
| GLM-4.7 Flash 30B-A3B (MoE) | 31.2B · 3B active | 19.8 GB | ≤ 110 | Fits |
| Gemma 4 31B | 31.3B | 20.9 GB | 11 | Fits |
| Nemotron 3.5 Lightning 30B-A3B (MoE) | 31.6B · 3B active | 19.7 GB | ≤ 110 | Fits |
| Qwen3 32B | 32.8B | 22.4 GB | 10 | Fits |
| Qwen2.5 Coder 32B | 32.8B | 22.4 GB | 10 | Fits |
| DeepSeek-R1 Distill Qwen 32B | 32.8B | 22.4 GB | 10 | Fits |
| Qwen3.6 35B-A3B (MoE) | 36B · 3B active | 22.4 GB | ≤ 110 | Fits |
| Llama 3.3 70B | 70.6B | 45.8 GB | 5 | Fits |
| DeepSeek-R1 Distill Llama 70B | 70.6B | 45.8 GB | 5 | Fits |
| Qwen3-Coder-Next 80B-A3B (MoE) | 79.7B · 3B active | 48.9 GB | ≤ 110 | Fits |
| gpt-oss 120B | 117B · 5.1B active | 71.4 GB | 65 | Fits |
| Mistral Small 4 119B (MoE) | 119B · 6.5B active | 72.5 GB | 51 | Fits |
On this card that means anything at or under 119.7 GB counts as fitting, and anything above 107.1 GB is marked tight: it loads, but there is little left for the conversation, and a desktop on the same 128 GB can push it over.
The model to actually run on it
Mistral Small 4 119B (MoE) is the best use of this card: 72.5 GB of the 126 GB available, an estimated 51 tokens per second, comfortable for chat, and 131,072 tokens of context still available. It is a mixture-of-experts model, so all 119B parameters sit in memory but only 6.5B are read per token — which is why it is quick for its size.
How this is chosen, since no benchmark is quoted: models are compared by size, a mixture-of-experts model counting at the geometric mean of its total and active parameters (a rule of thumb, not a measurement). The pick is the biggest class that fits with memory to spare and answers at 25 tokens per second or better, and within that class the most recent general-purpose release; coding and reasoning specialists are listed in the table but not recommended by default.
If quality matters more than parameter count, Llama 3.3 70B fits at Q8_0 in about 81 GB. Q8 is near-lossless; a smaller model at Q8 often beats a larger one squeezed into Q3.
How much context actually fits
Model size is the question people ask; context is the one that bites. The KV cache grows linearly with the conversation, so a model that loads comfortably can still run out of memory halfway through a long document. On 126 GB, the context you get is whatever the weights leave behind — which is why two cards that run the same model can feel completely different in use. These are the largest contexts this card holds, at Q4_K_M weights.
| Model | Max context, FP16 KV | With Q8 KV cache |
|---|---|---|
| Mistral Small 4 119B (MoE) | 131,072 tokens | 131,072 tokens |
| Llama 3.3 70B | 131,072 tokens | 131,072 tokens |
| Qwen3 32B | 40,960 tokens · model limit | 40,960 tokens · model limit |
| Gemma 4 31B | 131,072 tokens | 131,072 tokens |
Every model in this table already reaches its limit at FP16, so quantising the KV cache buys nothing here. Capped at 128k tokens, or at the model's own native window where that is smaller, marked "model limit": free memory beyond that point buys nothing. Rows that hold far more context than their size suggests are hybrid, sliding-window or latent-attention models, which cache only a fraction of what a standard transformer does; each model page shows the working.
Where this card stops
Nothing in this list is out of reach. Every model tracked here fits at Q4_K_M with 8k of context, so the ceiling you will meet is context length and batch size, not parameter count.
The honest take
It holds models no consumer card can, and it runs CUDA, which a Mac and an AMD box do not. The catch is 273 GB/s. A dense 70B model fits comfortably and crawls; mixture-of-experts models, which keep many parameters but read only a few billion per token, are what this machine is for. Think of it as a CUDA development box with a lot of memory, not a fast one.
What it is good at
- CUDA on 128 GB of unified memory: the NVIDIA software stack, from vLLM to TensorRT-LLM, runs as it does on a datacenter GPU.
- Blackwell, so FP4 weight formats run natively.
- A 150 mm box on a 240 W power supply, quiet enough for a desk.
What it is not
- 273 GB/s: dense models above 30B generate slowly however well they fit.
- An Arm Linux machine running DGX OS. Tools need ARM64 builds; Windows software does not run.
- nvidia-smi reports memory use as "Not Supported" and cudaMemGetInfo under-reports, which makes the memory look missing when it is not.
The thing people get wrong: NVIDIA's "up to 200 billion parameters" assumes 4-bit weights, as the datasheet's own footnote says. At Q8 the same memory holds roughly half that.
Buy or rent
Buy it for CUDA and capacity in a small box, not for speed: a used 3090 generates faster on anything that fits in 24 GB. If the large models are an occasional experiment, renting a large card for those hours costs less than a machine that runs them slowly every day.
Retail prices drift, and a figure written today would be wrong within a month, so this page does not carry one. Put the figure you are actually looking at, with your electricity rate and your real monthly hours, into the build vs rent calculator — it already knows this card draws 140 W under load. It returns the month owning becomes cheaper, or tells you it never does. Hours per month decides it far more often than the price does.
Nodegrove is the rented side of that comparison: a workspace that stays saved, with a GPU attached only while you are using it. It is the right answer when your hours are low or your needs change; buying is the right answer when they are high and stable.
Compared with the alternatives
- Ryzen AI Max+ 395 — 96 GB · 256 GB/s · best fit Mistral Small 4 119B (MoE) at ~48 tok/s. AMD's 128 GB laptop and mini-PC chip: up to 96 GB for the GPU on Windows, and an x86 route to large mixture-of-experts models.
- MacBook Pro M5 Max — 96 GB · 614 GB/s · best fit Qwen3.8 27B at ~27 tok/s. The laptop for large models in 2026: 128 GB at 614 GB/s, if you buy the 40-core GPU.
- RTX 5090 — 32 GB · 1,792 GB/s · best fit Qwen3.8 27B at ~78 tok/s. The first consumer card whose memory bandwidth changes what is comfortable.
The unified-memory machines are compared head to head, on the same models, in DGX Spark vs Strix Halo vs Mac Studio.
All 20 cards are compared side by side on the GPU index, and the same numbers from the model's point of view are on each model fit table. To check one specific pairing rather than read a table, use the can-I-run-it checker.
Questions
Can a DGX Spark run a 70B model?
Yes. Llama 3.3 70B at Q4_K_M with 8k of context needs about 45.8 GB, inside the 126 GB this card can address, and it generates an estimated 5 tokens per second — too slow for interactive use.
What is the best model to run on a DGX Spark?
Mistral Small 4 119B (MoE). At Q4_K_M and 8k context it needs about 72.5 GB of the 126 GB available and generates an estimated 51 tokens per second, which is comfortable for chat. If quality matters more than size, Llama 3.3 70B fits at Q8 in about 81 GB.
How many tokens per second does a DGX Spark generate?
It depends entirely on the size of the model, because generating a token means reading every active weight out of memory. At 273 GB/s and 70% efficiency, this card produces an estimated 41 tokens per second on an 8B model at Q4, and about 51 on the largest model it holds, Mistral Small 4 119B (MoE). Reading speed is roughly 5 to 8 words per second, so anything above 25 feels immediate. Treat 41 as a ceiling on 273 GB/s rather than a measurement; the higher a figure, the further real runtimes fall below it.
Is 126 GB enough for running LLMs locally?
It runs 29 of the 29 models tracked here at Q4_K_M with 8k of context, up to 119B parameters. The honest test is not the model list but the context: Mistral Small 4 119B (MoE) on this card holds about 131,072 tokens before memory runs out.
DGX Spark or Ryzen AI Max+ 395 for local models?
Both address about the same memory, so they run the same models. On speed, this card is faster: 273 GB/s against 256 GB/s, and bandwidth is what sets chat speed.
How much of the DGX Spark's 128 GB can the GPU use?
NVIDIA does not publish a figure. The GPU shares the whole pool with the Arm CPU and DGX OS, less a 2 GB display reserve (4 GB is an option in the BIOS), so what a model can use is whatever the system leaves free. This site plans around 126 GB and keeps a 5% margin for the operating system. Do not judge it by nvidia-smi, which shows memory use as "Not Supported" on this machine.
DGX Spark or Mac Studio for local LLMs?
The Spark runs CUDA, so the NVIDIA-only half of the ecosystem, including most fine-tuning and many image and video tools, works on it. A Mac Studio M5 Ultra reads memory at 1,200 GB/s against 273 GB/s and goes up to 512 GB, so it generates faster and holds more, without CUDA. For chatting with big models, the Mac; for building on NVIDIA hardware or fine-tuning, the Spark.
- Memory: weights (parameters × bytes per parameter) + KV cache (2 × layers × KV heads × head dim × context × bytes for a standard transformer; sliding-window, hybrid and latent-attention models counted as they actually cache) + 0.5 GB runtime + 4% of weights. Full derivation in the VRAM calculator.
- Speed: 70% × bandwidth ÷ active weight bytes, single stream, no batching, prompt processing excluded. A ceiling, not a measurement: the 6 of 29 figures marked ≤ are where real runtimes fall furthest below the number, because so few weights are read per token. See the tokens-per-second estimator.
- Card specification: 128 GB, 273 GB/s, 140 W — manufacturer figures; see NVIDIA's own page for this machine. Unified memory: NVIDIA publishes no GPU share: the GPU can use whatever of the 128 GB the operating system leaves free, less a 2 GB display reserve. Figures here use 126 GB; the operating system comes out of the 5% margin every fit keeps, so run the largest models without a desktop session. Source: NVIDIA's DGX Spark release notes (the 2 GB display reserve).
- Model architecture values come from each model's published config; the self-hosted LLM guide explains what each one changes.