All three put a large pool of memory beside the GPU, which is what lets a small box hold models no gaming card can. They differ on the two things that decide everything else: how fast the GPU reads that memory, and which software stack it runs. The Mac Studio M5 Ultra reads memory at 1,200 GB/s, more than 4 times the DGX Spark's 273 GB/s and the Ryzen AI Max+ 395's 256 GB/s. The DGX Spark is the only one that runs CUDA. The Ryzen AI Max+ 395 is the only one that runs Windows.
The numbers that matter
| Machine | Memory | GPU can use | Bandwidth | Software |
|---|---|---|---|---|
| DGX Spark | 128 GB | 126 GB | 273 GB/s | CUDA on Arm Linux (DGX OS) |
| Ryzen AI Max+ 395 | 128 GB | 96 GB | 256 GB/s | ROCm or Vulkan on Windows or Linux |
| MacBook Pro M5 Max | 128 GB | 96 GB | 614 GB/s | Metal and MLX on macOS |
| Mac Studio M3 Ultra | 512 GB | 384 GB | 819 GB/s | Metal and MLX on macOS |
| Mac Studio M5 Ultra | 512 GB | 384 GB | 1,200 GB/s | Metal and MLX on macOS |
"GPU can use" is the figure every fit on this site is measured against. DGX Spark: NVIDIA publishes no GPU share: the GPU can use whatever of the 128 GB the operating system leaves free, less a 2 GB display reserve. Figures here use 126 GB; the operating system comes out of the 5% margin every fit keeps, so run the largest models without a desktop session. Ryzen AI Max+ 395: AMD lets up to 96 GB of the 128 GB be set aside as graphics memory (Variable Graphics Memory), and says workloads run best inside it. Figures here use 96 GB. On Linux AMD documents raising the shared limit further, to 120 GB in its own guide. On the Macs it is about 75% of RAM, the macOS limit observed on large-memory Macs; Apple publishes no figure. Each machine is shown at its largest memory option. The 512 GB M5 Ultra needs the top 80-core GPU chip and ships in late October 2026; the 512 GB M3 Ultra was offered at launch and is no longer on Apple's spec page.
For scale: an RTX 4090 reads its 24 GB at 1,008 GB/s. The DGX Spark and the Ryzen AI Max+ 395 hold far more and read it at about a quarter of that speed.
What each one runs
Of the 29 open models tracked on this site, at Q4_K_M with 8k of context:
| Machine | Runs | Largest | The one to run |
|---|---|---|---|
| DGX Spark | 29 of 29 | Mistral Small 4 119B (MoE) | Mistral Small 4 119B (MoE) 72.5 GB |
| Ryzen AI Max+ 395 | 29 of 29 | Mistral Small 4 119B (MoE) | Mistral Small 4 119B (MoE) 72.5 GB |
| MacBook Pro M5 Max | 29 of 29 | Mistral Small 4 119B (MoE) | Qwen3.8 27B 18 GB |
| Mac Studio M3 Ultra | 29 of 29 | Mistral Small 4 119B (MoE) | Qwen3.8 27B 18 GB |
| Mac Studio M5 Ultra | 29 of 29 | Mistral Small 4 119B (MoE) | Qwen3.8 27B 18 GB |
"The one to run" is the same pick as on each GPU page: the biggest class that fits with room for a conversation and generates at a usable speed, then the newest release. Capacity stops being the difference here: every machine in this comparison runs all 29 of them, so the choice comes down to speed and software.
Speed on the same models
Estimated tokens per second at Q4_K_M and 8k of context, single stream. These are ceilings worked out from memory bandwidth, not measurements; figures marked ≤ are where real runtimes fall furthest below.
| Model | DGX Spark | Ryzen AI Max+ 395 | MacBook Pro M5 Max | Mac Studio M3 Ultra | Mac Studio M5 Ultra |
|---|---|---|---|---|---|
| Qwen3.8 27B | 12 | 11 | 27 | 36 | 52 |
| gpt-oss 120B | 65 | 61 | ≤ 145 | ≤ 194 | ≤ 284 |
| Qwen3-Coder-Next 80B-A3B (MoE) | ≤ 110 | ≤ 103 | ≤ 247 | ≤ 329 | ≤ 483 |
| Mistral Small 4 119B (MoE) | 51 | 48 | ≤ 114 | ≤ 152 | ≤ 223 |
| Llama 3.3 70B | 5 | 4 | 10 | 14 | 21 |
Mixture-of-experts models are where the DGX Spark and the Ryzen AI Max+ 395 make sense: gpt-oss 120B reads only a few billion parameters per token, so it runs at about 65 tokens per second on the Spark and about 61 on the AMD chip, while a dense 70B model manages about 5 on the Spark. These figures describe generation; reading a long prompt first is a separate cost, and it is markedly slower on Apple silicon than on NVIDIA hardware.
Which to choose
- DGX Spark if you need CUDA: fine-tuning, TensorRT-LLM, vLLM as on a datacenter card, or image and video tools that ship CUDA builds. It is an Arm Linux box, not a desktop for everything else.
- Ryzen AI Max+ 395 if you want one x86 machine that also holds large mixture-of-experts models, including a laptop. Check the power limit of the specific machine: makers set it between 45 and 120 W.
- A Mac Studio if you want the most memory and the fastest generation in a quiet box and do not need CUDA. The M5 Ultra is the faster one; a 512 GB M3 Ultra holds the same models. The M5 Max is the same idea at 128 GB.
None of them is the cheap way to run a model now and then. If the large models are an occasional need, renting a large card for those hours costs less than a machine sized for them; the build vs rent calculator finds the break-even with your own numbers.
Questions
DGX Spark or Strix Halo for local LLMs?
They are close on speed: 273 GB/s against 256 GB/s. The Spark gives the GPU more memory (126 GB against 96 GB on Windows) and runs CUDA, so the NVIDIA software stack works as it does on a datacenter card. The Ryzen AI Max+ 395 is an x86 chip in laptops and mini-PCs that runs Windows; it relies on ROCm or Vulkan. Both run gpt-oss 120B at Q4 (71.4 GB); for CUDA tooling, the Spark, for a Windows machine, the AMD.
DGX Spark or Mac Studio for local LLMs?
A Mac Studio M5 Ultra reads memory about 4.4 times as fast (1,200 GB/s against 273) and gives the GPU up to 384 GB, so it generates faster and holds more: Llama 3.3 70B at Q4 runs at about 21 tokens per second against about 5 on the Spark. The Spark runs CUDA, which most fine-tuning and much image and video tooling needs, and processes long prompts faster than Apple silicon.
Is Strix Halo or a Mac Studio better for local LLMs?
For speed and capacity, the Mac Studio: even the M3 Ultra reads memory 3.2 times as fast as the Ryzen AI Max+ 395. The AMD chip's case is an x86 machine that runs Windows and Linux, from several makers, in a laptop if you want one.
How much memory can the GPU use on each?
DGX Spark: NVIDIA publishes no share; the GPU uses what DGX OS leaves free, less a 2 GB display reserve, so this site plans around 126 GB. Ryzen AI Max+ 395: up to 96 GB on Windows with AMD Variable Graphics Memory, more on Linux. Macs: about 75% of RAM by default, an observed macOS limit that Apple does not publish.
- Memory and bandwidth: each maker's specification, linked on each machine's GPU page, with bandwidth sources where the spec page does not print the figure.
- GPU share of memory: NVIDIA's DGX Spark release notes (2 GB display reserve), AMD's Variable Graphics Memory FAQ (96 GB), and the macOS limit observed on Apple silicon.
- Fits and speeds: the same formulas as the VRAM calculator and the speed estimator, on each model's config.json.