128 GB unified 614 GB/s 2026 Apple unified memory
MacBook Pro M5 Max (128 GB) addresses 96 GB — 75% of its 128 GB of unified memory, the share macOS is observed to hand the GPU by default, at 614 GB/s. Of the 29 open models tracked on this site, it runs 29 at Q4_K_M with 8k of context. The largest is Mistral Small 4 119B (MoE), needing about 72.5 GB and generating an estimated 114 tokens per second — faster than you can read.
The laptop for large models in 2026: 128 GB at 614 GB/s, if you buy the 40-core GPU.
What a MacBook Pro M5 Max runs, model by model
Every model at Q4_K_M with 8k of context, the setting most people actually use. This card has 91.2 GB to spend once the 5% safety margin comes off its 96 GB, and it reads that memory at 614 GB/s. Those two numbers decide the last two columns: the first says what loads, the second says how fast it answers. Both are estimates from stated formulas, and the VRAM calculator shows the working.
| Model | Size | Needs | Est. tokens/s | On this card |
|---|---|---|---|---|
| Llama 3.2 3B | 3.2B | 3.4 GB | ≤ 232 | Fits |
| Qwen3.5 4B | 4.7B | 3.7 GB | ≤ 158 | Fits |
| Llama 3.1 8B | 8B | 6.4 GB | 93 | Fits |
| Qwen3 8B | 8.2B | 6.7 GB | 90 | Fits |
| Qwen3.5 9B | 9.7B | 6.7 GB | 76 | Fits |
| Gemma 4 12B | 12B | 8.2 GB | 62 | Fits |
| Gemma 3 12B | 12.2B | 8.7 GB | 61 | Fits |
| Ministral 3 14B | 14B | 10.3 GB | 53 | Fits |
| Qwen3 14B | 14.8B | 10.8 GB | 50 | Fits |
| Phi-4 14B | 14.7B | 11 GB | 50 | Fits |
| gpt-oss 20B | 21B · 3.6B active | 13.4 GB | ≤ 206 | Fits |
| Devstral Small 2 24B | 24B | 16.3 GB | 31 | Fits |
| Mistral Small 3.1 24B | 24B | 16.3 GB | 31 | Fits |
| Gemma 4 26B-A4B (MoE) | 25.8B · 3.8B active | 16.4 GB | ≤ 195 | Fits |
| Gemma 3 27B | 27.4B | 18.1 GB | 27 | Fits |
| Qwen3.8 27B | 27.8B | 18 GB | 27 | Fits |
| Qwen3 30B-A3B (MoE) | 30.5B · 3.3B active | 19.7 GB | ≤ 225 | Fits |
| GLM-4.7 Flash 30B-A3B (MoE) | 31.2B · 3B active | 19.8 GB | ≤ 247 | Fits |
| Gemma 4 31B | 31.3B | 20.9 GB | 24 | Fits |
| Nemotron 3.5 Lightning 30B-A3B (MoE) | 31.6B · 3B active | 19.7 GB | ≤ 247 | Fits |
| Qwen3 32B | 32.8B | 22.4 GB | 23 | Fits |
| Qwen2.5 Coder 32B | 32.8B | 22.4 GB | 23 | Fits |
| DeepSeek-R1 Distill Qwen 32B | 32.8B | 22.4 GB | 23 | Fits |
| Qwen3.6 35B-A3B (MoE) | 36B · 3B active | 22.4 GB | ≤ 247 | Fits |
| Llama 3.3 70B | 70.6B | 45.8 GB | 10 | Fits |
| DeepSeek-R1 Distill Llama 70B | 70.6B | 45.8 GB | 10 | Fits |
| Qwen3-Coder-Next 80B-A3B (MoE) | 79.7B · 3B active | 48.9 GB | ≤ 247 | Fits |
| gpt-oss 120B | 117B · 5.1B active | 71.4 GB | ≤ 145 | Fits |
| Mistral Small 4 119B (MoE) | 119B · 6.5B active | 72.5 GB | ≤ 114 | Fits |
On this card that means anything at or under 91.2 GB counts as fitting, and anything above 81.6 GB is marked tight: it loads, but there is little left for the conversation, and a desktop on the same 128 GB can push it over.
One caveat specific to Apple silicon: these tokens-per-second figures describe generation, which unified memory handles well. Prompt processing — reading what you send before the first token comes back — is markedly slower than on an equivalent Nvidia card, and that is what you feel when you paste in a long document. Short prompts feel fast on this machine; 30,000-token ones do not.
The model to actually run on it
Qwen3.8 27B is the best use of this card: 18 GB of the 96 GB available, an estimated 27 tokens per second, comfortable for chat, and 131,072 tokens of context still available.
How this is chosen, since no benchmark is quoted: models are compared by size, a mixture-of-experts model counting at the geometric mean of its total and active parameters (a rule of thumb, not a measurement). The pick is the biggest class that fits with memory to spare and answers at 25 tokens per second or better, and within that class the most recent general-purpose release; coding and reasoning specialists are listed in the table but not recommended by default.
Mistral Small 4 119B (MoE) is the largest model the card will hold, at 72.5 GB, with room for 131,072 tokens of context. It is a mixture-of-experts model reading only 6.5B parameters per token, so it is the faster of the two at an estimated 114 tokens per second; what it gives up is the depth of a dense model that reads all of its weights for every token.
If quality matters more than parameter count, Llama 3.3 70B fits at Q8_0 in about 81 GB. Q8 is near-lossless; a smaller model at Q8 often beats a larger one squeezed into Q3.
How much context actually fits
Model size is the question people ask; context is the one that bites. The KV cache grows linearly with the conversation, so a model that loads comfortably can still run out of memory halfway through a long document. On 96 GB, the context you get is whatever the weights leave behind — which is why two cards that run the same model can feel completely different in use. These are the largest contexts this card holds, at Q4_K_M weights.
| Model | Max context, FP16 KV | With Q8 KV cache |
|---|---|---|
| Qwen3.8 27B | 131,072 tokens | 131,072 tokens |
| Llama 3.3 70B | 131,072 tokens | 131,072 tokens |
| Qwen3 32B | 40,960 tokens · model limit | 40,960 tokens · model limit |
| Gemma 4 31B | 131,072 tokens | 131,072 tokens |
Every model in this table already reaches its limit at FP16, so quantising the KV cache buys nothing here. Capped at 128k tokens, or at the model's own native window where that is smaller, marked "model limit": free memory beyond that point buys nothing. Rows that hold far more context than their size suggests are hybrid, sliding-window or latent-attention models, which cache only a fraction of what a standard transformer does; each model page shows the working.
Where this card stops
Nothing in this list is out of reach. Every model tracked here fits at Q4_K_M with 8k of context, so the ceiling you will meet is context length and batch size, not parameter count.
The honest take
The M4 Max's ceiling of 128 GB with faster memory: 614 GB/s against 546 GB/s. That holds a 70B model at Q4 with a long context, on battery, at a speed that is fine for chat. Prompt processing is slower than on Nvidia, as on every Apple machine, so long documents pause before the first token. The cheaper M5 Max with the 32-core GPU is a different machine for this purpose: 36 GB and less bandwidth.
What it is good at
- Up to 128 GB of unified memory in a laptop, with about 96 GB for the GPU.
- 614 GB/s, about an eighth faster than the M4 Max.
- MLX, llama.cpp, Ollama and LM Studio all run on it.
What it is not
- Prompt processing lags Nvidia badly, which is felt most on long documents and agent loops.
- No CUDA: much of the image, video and fine-tuning ecosystem is unavailable or slow.
- Memory is fixed at purchase.
The thing people get wrong: Only the 40-core GPU version offers 48, 64 or 128 GB and the full 614 GB/s. The 32-core GPU M5 Max comes with 36 GB only.
Buy or rent
If you were buying a MacBook Pro anyway, the 128 GB configuration is the most capable portable machine for local models. If you were not, it is an expensive way to buy GPU memory, and it does not run CUDA.
What matters is the memory upgrade rather than the machine, and Apple prices that differently in every configuration. Put the figure you are actually looking at, with your electricity rate and your real monthly hours, into the build vs rent calculator. It returns the month owning becomes cheaper, or tells you it never does. Hours per month decides it far more often than the price does.
Nodegrove is the rented side of that comparison: a workspace that stays saved, with a GPU attached only while you are using it. It is the right answer when your hours are low or your needs change; buying is the right answer when they are high and stable.
Compared with the alternatives
- MacBook Pro M4 Max — 96 GB · 546 GB/s · best fit Mistral Small 4 119B (MoE) at ~101 tok/s. A laptop that runs models a desktop GPU cannot hold.
- Mac Studio M5 Ultra — 384 GB · 1,200 GB/s · best fit Qwen3.8 27B at ~52 tok/s. The most memory you can put on a desk: up to 512 GB at 1.2 TB/s, without CUDA.
- DGX Spark — 126 GB · 273 GB/s · best fit Mistral Small 4 119B (MoE) at ~51 tok/s. NVIDIA's small Linux box with 128 GB the GPU can share: large models fit, but they are read at about a quarter of an RTX 4090's speed.
The unified-memory machines are compared head to head, on the same models, in DGX Spark vs Strix Halo vs Mac Studio.
All 20 cards are compared side by side on the GPU index, and the same numbers from the model's point of view are on each model fit table. To check one specific pairing rather than read a table, use the can-I-run-it checker.
Questions
Can a MacBook Pro M5 Max run a 70B model?
Yes. Llama 3.3 70B at Q4_K_M with 8k of context needs about 45.8 GB, inside the 96 GB this card can address, and it generates an estimated 10 tokens per second — slow; fine for batch jobs.
What is the best model to run on a MacBook Pro M5 Max?
Qwen3.8 27B. At Q4_K_M and 8k context it needs about 18 GB of the 96 GB available and generates an estimated 27 tokens per second, which is comfortable for chat. If quality matters more than size, Llama 3.3 70B fits at Q8 in about 81 GB.
How many tokens per second does a MacBook Pro M5 Max generate?
It depends entirely on the size of the model, because generating a token means reading every active weight out of memory. At 614 GB/s and 70% efficiency, this card produces an estimated 93 tokens per second on an 8B model at Q4, and about 114 on the largest model it holds, Mistral Small 4 119B (MoE). Reading speed is roughly 5 to 8 words per second, so anything above 25 feels immediate. Treat 93 as a ceiling on 614 GB/s rather than a measurement; the higher a figure, the further real runtimes fall below it.
Is 96 GB enough for running LLMs locally?
It runs 29 of the 29 models tracked here at Q4_K_M with 8k of context, up to 119B parameters. The honest test is not the model list but the context: Qwen3.8 27B on this card holds about 131,072 tokens before memory runs out.
MacBook Pro M5 Max or MacBook Pro M4 Max for local models?
Both address about the same memory, so they run the same models. On speed, this card is faster: 614 GB/s against 546 GB/s, and bandwidth is what sets chat speed.
M5 Max or M4 Max for local LLMs?
The same 128 GB ceiling, so the same models fit. The M5 Max reads memory at 614 GB/s against 546 GB/s, so it generates text about an eighth faster. If you already own a 128 GB M4 Max, that is not a reason to upgrade.
Which M5 Max should I buy for AI?
The 40-core GPU version. It is the only one offered with 48, 64 or 128 GB and the full 614 GB/s; the 32-core GPU version comes with 36 GB only. The GPU gets about 75% of whatever you buy, and memory cannot be upgraded later.
- Memory: weights (parameters × bytes per parameter) + KV cache (2 × layers × KV heads × head dim × context × bytes for a standard transformer; sliding-window, hybrid and latent-attention models counted as they actually cache) + 0.5 GB runtime + 4% of weights. Full derivation in the VRAM calculator.
- Speed: 70% × bandwidth ÷ active weight bytes, single stream, no batching, prompt processing excluded. A ceiling, not a measurement: the 11 of 29 figures marked ≤ are where real runtimes fall furthest below the number, because so few weights are read per token. See the tokens-per-second estimator.
- Card specification: 128 GB, 614 GB/s — manufacturer figures; see Apple's own page for this machine. Apple unified memory: Apple publishes no GPU share; macOS's Metal limit is observed at about 75% of RAM on large-memory Macs, and every figure here uses that share.
- Model architecture values come from each model's published config; the self-hosted LLM guide explains what each one changes.