512 GB unified 1,200 GB/s 2026 Apple unified memory
Mac Studio M5 Ultra (512 GB) addresses 384 GB — 75% of its 512 GB of unified memory, the share macOS is observed to hand the GPU by default, at 1,200 GB/s. Of the 29 open models tracked on this site, it runs 29 at Q4_K_M with 8k of context. The largest is Mistral Small 4 119B (MoE), needing about 72.5 GB and generating an estimated 223 tokens per second — faster than you can read.
The most memory you can put on a desk: up to 512 GB at 1.2 TB/s, without CUDA.
What a Mac Studio M5 Ultra runs, model by model
Every model at Q4_K_M with 8k of context, the setting most people actually use. This card has 364.8 GB to spend once the 5% safety margin comes off its 384 GB, and it reads that memory at 1,200 GB/s. Those two numbers decide the last two columns: the first says what loads, the second says how fast it answers. Both are estimates from stated formulas, and the VRAM calculator shows the working.
| Model | Size | Needs | Est. tokens/s | On this card |
|---|---|---|---|---|
| Llama 3.2 3B | 3.2B | 3.4 GB | ≤ 453 | Fits |
| Qwen3.5 4B | 4.7B | 3.7 GB | ≤ 308 | Fits |
| Llama 3.1 8B | 8B | 6.4 GB | ≤ 181 | Fits |
| Qwen3 8B | 8.2B | 6.7 GB | ≤ 177 | Fits |
| Qwen3.5 9B | 9.7B | 6.7 GB | ≤ 149 | Fits |
| Gemma 4 12B | 12B | 8.2 GB | ≤ 121 | Fits |
| Gemma 3 12B | 12.2B | 8.7 GB | ≤ 119 | Fits |
| Ministral 3 14B | 14B | 10.3 GB | ≤ 103 | Fits |
| Qwen3 14B | 14.8B | 10.8 GB | 98 | Fits |
| Phi-4 14B | 14.7B | 11 GB | 99 | Fits |
| gpt-oss 20B | 21B · 3.6B active | 13.4 GB | ≤ 402 | Fits |
| Devstral Small 2 24B | 24B | 16.3 GB | 60 | Fits |
| Mistral Small 3.1 24B | 24B | 16.3 GB | 60 | Fits |
| Gemma 4 26B-A4B (MoE) | 25.8B · 3.8B active | 16.4 GB | ≤ 381 | Fits |
| Gemma 3 27B | 27.4B | 18.1 GB | 53 | Fits |
| Qwen3.8 27B | 27.8B | 18 GB | 52 | Fits |
| Qwen3 30B-A3B (MoE) | 30.5B · 3.3B active | 19.7 GB | ≤ 439 | Fits |
| GLM-4.7 Flash 30B-A3B (MoE) | 31.2B · 3B active | 19.8 GB | ≤ 483 | Fits |
| Gemma 4 31B | 31.3B | 20.9 GB | 46 | Fits |
| Nemotron 3.5 Lightning 30B-A3B (MoE) | 31.6B · 3B active | 19.7 GB | ≤ 483 | Fits |
| Qwen3 32B | 32.8B | 22.4 GB | 44 | Fits |
| Qwen2.5 Coder 32B | 32.8B | 22.4 GB | 44 | Fits |
| DeepSeek-R1 Distill Qwen 32B | 32.8B | 22.4 GB | 44 | Fits |
| Qwen3.6 35B-A3B (MoE) | 36B · 3B active | 22.4 GB | ≤ 483 | Fits |
| Llama 3.3 70B | 70.6B | 45.8 GB | 21 | Fits |
| DeepSeek-R1 Distill Llama 70B | 70.6B | 45.8 GB | 21 | Fits |
| Qwen3-Coder-Next 80B-A3B (MoE) | 79.7B · 3B active | 48.9 GB | ≤ 483 | Fits |
| gpt-oss 120B | 117B · 5.1B active | 71.4 GB | ≤ 284 | Fits |
| Mistral Small 4 119B (MoE) | 119B · 6.5B active | 72.5 GB | ≤ 223 | Fits |
On this card that means anything at or under 364.8 GB counts as fitting, and anything above 326.4 GB is marked tight: it loads, but there is little left for the conversation, and a desktop on the same 512 GB can push it over.
One caveat specific to Apple silicon: these tokens-per-second figures describe generation, which unified memory handles well. Prompt processing — reading what you send before the first token comes back — is markedly slower than on an equivalent Nvidia card, and that is what you feel when you paste in a long document. Short prompts feel fast on this machine; 30,000-token ones do not.
The model to actually run on it
Qwen3.8 27B is the best use of this card: 18 GB of the 384 GB available, an estimated 52 tokens per second, comfortable for chat, and 131,072 tokens of context still available.
How this is chosen, since no benchmark is quoted: models are compared by size, a mixture-of-experts model counting at the geometric mean of its total and active parameters (a rule of thumb, not a measurement). The pick is the biggest class that fits with memory to spare and answers at 25 tokens per second or better, and within that class the most recent general-purpose release; coding and reasoning specialists are listed in the table but not recommended by default.
Mistral Small 4 119B (MoE) is the largest model the card will hold, at 72.5 GB, with room for 131,072 tokens of context. It is a mixture-of-experts model reading only 6.5B parameters per token, so it is the faster of the two at an estimated 223 tokens per second; what it gives up is the depth of a dense model that reads all of its weights for every token.
If quality matters more than parameter count, Llama 3.3 70B fits at Q8_0 in about 81 GB. Q8 is near-lossless; a smaller model at Q8 often beats a larger one squeezed into Q3.
How much context actually fits
Model size is the question people ask; context is the one that bites. The KV cache grows linearly with the conversation, so a model that loads comfortably can still run out of memory halfway through a long document. On 384 GB, the context you get is whatever the weights leave behind — which is why two cards that run the same model can feel completely different in use. These are the largest contexts this card holds, at Q4_K_M weights.
| Model | Max context, FP16 KV | With Q8 KV cache |
|---|---|---|
| Qwen3.8 27B | 131,072 tokens | 131,072 tokens |
| Llama 3.3 70B | 131,072 tokens | 131,072 tokens |
| Qwen3 32B | 40,960 tokens · model limit | 40,960 tokens · model limit |
| Gemma 4 31B | 131,072 tokens | 131,072 tokens |
Every model in this table already reaches its limit at FP16, so quantising the KV cache buys nothing here. Capped at 128k tokens, or at the model's own native window where that is smaller, marked "model limit": free memory beyond that point buys nothing. Rows that hold far more context than their size suggests are hybrid, sliding-window or latent-attention models, which cache only a fraction of what a standard transformer does; each model page shows the working.
Where this card stops
Nothing in this list is out of reach. Every model tracked here fits at Q4_K_M with 8k of context, so the ceiling you will meet is context length and batch size, not parameter count.
The honest take
A 512 GB M5 Ultra gives the GPU about 384 GB, enough for every model tracked on this site at Q8 and for mixture-of-experts models far larger than any single graphics card holds. It reads memory at 1,200 GB/s, faster than a 4090 and slower than a 5090. The usual Apple trade applies: generation is quick, prompt processing is not, and CUDA tooling does not run.
What it is good at
- Up to 512 GB of unified memory, about 384 GB of it for the GPU.
- 1,200 GB/s, about half as fast again as the M3 Ultra.
- Quiet, and a fraction of the power of a multi-GPU build.
What it is not
- Prompt processing is far slower than on Nvidia: long documents and agent loops feel it.
- No CUDA, which rules out much of the fine-tuning and image and video ecosystem.
- The 512 GB configuration needs the top 80-core GPU chip and ships in late October 2026.
The thing people get wrong: Apple publishes no GPU share of unified memory. macOS is observed to give the GPU about 75% by default, so a 512 GB machine plans around 384 GB. The limit can be raised with the iogpu.wired_limit_mb setting, at the risk of starving the system.
Buy or rent
The machine to buy if you want very large models locally, silently, and do not need CUDA. If you only need such models now and then, renting a large card for those hours costs far less than a machine sized for the largest one.
What matters is the memory upgrade rather than the machine, and Apple prices that differently in every configuration. Put the figure you are actually looking at, with your electricity rate and your real monthly hours, into the build vs rent calculator. It returns the month owning becomes cheaper, or tells you it never does. Hours per month decides it far more often than the price does.
Nodegrove is the rented side of that comparison: a workspace that stays saved, with a GPU attached only while you are using it. It is the right answer when your hours are low or your needs change; buying is the right answer when they are high and stable.
Compared with the alternatives
- Mac Studio M3 Ultra — 384 GB · 819 GB/s · best fit Qwen3.8 27B at ~36 tok/s. 2025's 512 GB Mac Studio: still the way to hold very large models, now mostly second-hand.
- DGX Spark — 126 GB · 273 GB/s · best fit Mistral Small 4 119B (MoE) at ~51 tok/s. NVIDIA's small Linux box with 128 GB the GPU can share: large models fit, but they are read at about a quarter of an RTX 4090's speed.
- RTX PRO 6000 Blackwell — 96 GB · 1,792 GB/s · best fit Llama 3.3 70B at ~31 tok/s. Ninety-six gigabytes and 5090-class bandwidth on one card.
The unified-memory machines are compared head to head, on the same models, in DGX Spark vs Strix Halo vs Mac Studio.
All 20 cards are compared side by side on the GPU index, and the same numbers from the model's point of view are on each model fit table. To check one specific pairing rather than read a table, use the can-I-run-it checker.
Questions
Can a Mac Studio M5 Ultra run a 70B model?
Yes. Llama 3.3 70B at Q4_K_M with 8k of context needs about 45.8 GB, inside the 384 GB this card can address, and it generates an estimated 21 tokens per second — usable, a little slow.
What is the best model to run on a Mac Studio M5 Ultra?
Qwen3.8 27B. At Q4_K_M and 8k context it needs about 18 GB of the 384 GB available and generates an estimated 52 tokens per second, which is comfortable for chat. If quality matters more than size, Llama 3.3 70B fits at Q8 in about 81 GB.
How many tokens per second does a Mac Studio M5 Ultra generate?
It depends entirely on the size of the model, because generating a token means reading every active weight out of memory. At 1,200 GB/s and 70% efficiency, this card produces an estimated 181 tokens per second on an 8B model at Q4, and about 223 on the largest model it holds, Mistral Small 4 119B (MoE). Reading speed is roughly 5 to 8 words per second, so anything above 25 feels immediate. Treat 181 as a ceiling on 1,200 GB/s rather than a measurement; the higher a figure, the further real runtimes fall below it.
Is 384 GB enough for running LLMs locally?
It runs 29 of the 29 models tracked here at Q4_K_M with 8k of context, up to 119B parameters. The honest test is not the model list but the context: Qwen3.8 27B on this card holds about 131,072 tokens before memory runs out.
Mac Studio M5 Ultra or Mac Studio M3 Ultra for local models?
Both address about the same memory, so they run the same models. On speed, this card is faster: 1,200 GB/s against 819 GB/s, and bandwidth is what sets chat speed.
M5 Ultra or M3 Ultra for local LLMs?
Both reach 512 GB, so the same models fit. The M5 Ultra reads memory at 1,200 GB/s against 819 GB/s, so it generates text about half as fast again. A second-hand 512 GB M3 Ultra is still a very capable machine; the M5 Ultra is the faster one.
How much of a 512 GB Mac Studio can the GPU use?
About 75% by default, so roughly 384 GB. Apple does not publish a figure; that is the macOS limit observed on large-memory Macs, and every figure on this page uses it.
- Memory: weights (parameters × bytes per parameter) + KV cache (2 × layers × KV heads × head dim × context × bytes for a standard transformer; sliding-window, hybrid and latent-attention models counted as they actually cache) + 0.5 GB runtime + 4% of weights. Full derivation in the VRAM calculator.
- Speed: 70% × bandwidth ÷ active weight bytes, single stream, no batching, prompt processing excluded. A ceiling, not a measurement: the 17 of 29 figures marked ≤ are where real runtimes fall furthest below the number, because so few weights are read per token. See the tokens-per-second estimator.
- Card specification: 512 GB, 1,200 GB/s — manufacturer figures; see Apple's own page for this machine. Apple unified memory: Apple publishes no GPU share; macOS's Metal limit is observed at about 75% of RAM on large-memory Macs, and every figure here uses that share.
- Model architecture values come from each model's published config; the self-hosted LLM guide explains what each one changes.