What LLMs can a MacBook Pro M5 Max run?

Which open models run on a MacBook Pro M5 Max: what fits at Q4 and Q8, estimated tokens per second, how much context fits, and where the card stops.

Updated · estimates are labelled as estimates
On this page

128 GB unified 614 GB/s 2026 Apple unified memory

MacBook Pro M5 Max (128 GB) addresses 96 GB — 75% of its 128 GB of unified memory, the share macOS is observed to hand the GPU by default, at 614 GB/s. Of the 29 open models tracked on this site, it runs 29 at Q4_K_M with 8k of context. The largest is Mistral Small 4 119B (MoE), needing about 72.5 GB and generating an estimated 114 tokens per second — faster than you can read.

The laptop for large models in 2026: 128 GB at 614 GB/s, if you buy the 40-core GPU.

What a MacBook Pro M5 Max runs, model by model

Every model at Q4_K_M with 8k of context, the setting most people actually use. This card has 91.2 GB to spend once the 5% safety margin comes off its 96 GB, and it reads that memory at 614 GB/s. Those two numbers decide the last two columns: the first says what loads, the second says how fast it answers. Both are estimates from stated formulas, and the VRAM calculator shows the working.

ModelSizeNeedsEst. tokens/sOn this card
Llama 3.2 3B 3.2B 3.4 GB ≤ 232 Fits
Qwen3.5 4B 4.7B 3.7 GB ≤ 158 Fits
Llama 3.1 8B 8B 6.4 GB 93 Fits
Qwen3 8B 8.2B 6.7 GB 90 Fits
Qwen3.5 9B 9.7B 6.7 GB 76 Fits
Gemma 4 12B 12B 8.2 GB 62 Fits
Gemma 3 12B 12.2B 8.7 GB 61 Fits
Ministral 3 14B 14B 10.3 GB 53 Fits
Qwen3 14B 14.8B 10.8 GB 50 Fits
Phi-4 14B 14.7B 11 GB 50 Fits
gpt-oss 20B 21B · 3.6B active 13.4 GB ≤ 206 Fits
Devstral Small 2 24B 24B 16.3 GB 31 Fits
Mistral Small 3.1 24B 24B 16.3 GB 31 Fits
Gemma 4 26B-A4B (MoE) 25.8B · 3.8B active 16.4 GB ≤ 195 Fits
Gemma 3 27B 27.4B 18.1 GB 27 Fits
Qwen3.8 27B 27.8B 18 GB 27 Fits
Qwen3 30B-A3B (MoE) 30.5B · 3.3B active 19.7 GB ≤ 225 Fits
GLM-4.7 Flash 30B-A3B (MoE) 31.2B · 3B active 19.8 GB ≤ 247 Fits
Gemma 4 31B 31.3B 20.9 GB 24 Fits
Nemotron 3.5 Lightning 30B-A3B (MoE) 31.6B · 3B active 19.7 GB ≤ 247 Fits
Qwen3 32B 32.8B 22.4 GB 23 Fits
Qwen2.5 Coder 32B 32.8B 22.4 GB 23 Fits
DeepSeek-R1 Distill Qwen 32B 32.8B 22.4 GB 23 Fits
Qwen3.6 35B-A3B (MoE) 36B · 3B active 22.4 GB ≤ 247 Fits
Llama 3.3 70B 70.6B 45.8 GB 10 Fits
DeepSeek-R1 Distill Llama 70B 70.6B 45.8 GB 10 Fits
Qwen3-Coder-Next 80B-A3B (MoE) 79.7B · 3B active 48.9 GB ≤ 247 Fits
gpt-oss 120B 117B · 5.1B active 71.4 GB ≤ 145 Fits
Mistral Small 4 119B (MoE) 119B · 6.5B active 72.5 GB ≤ 114 Fits

On this card that means anything at or under 91.2 GB counts as fitting, and anything above 81.6 GB is marked tight: it loads, but there is little left for the conversation, and a desktop on the same 128 GB can push it over.

One caveat specific to Apple silicon: these tokens-per-second figures describe generation, which unified memory handles well. Prompt processing — reading what you send before the first token comes back — is markedly slower than on an equivalent Nvidia card, and that is what you feel when you paste in a long document. Short prompts feel fast on this machine; 30,000-token ones do not.

The model to actually run on it

Qwen3.8 27B is the best use of this card: 18 GB of the 96 GB available, an estimated 27 tokens per second, comfortable for chat, and 131,072 tokens of context still available.

How this is chosen, since no benchmark is quoted: models are compared by size, a mixture-of-experts model counting at the geometric mean of its total and active parameters (a rule of thumb, not a measurement). The pick is the biggest class that fits with memory to spare and answers at 25 tokens per second or better, and within that class the most recent general-purpose release; coding and reasoning specialists are listed in the table but not recommended by default.

Mistral Small 4 119B (MoE) is the largest model the card will hold, at 72.5 GB, with room for 131,072 tokens of context. It is a mixture-of-experts model reading only 6.5B parameters per token, so it is the faster of the two at an estimated 114 tokens per second; what it gives up is the depth of a dense model that reads all of its weights for every token.

If quality matters more than parameter count, Llama 3.3 70B fits at Q8_0 in about 81 GB. Q8 is near-lossless; a smaller model at Q8 often beats a larger one squeezed into Q3.

How much context actually fits

Model size is the question people ask; context is the one that bites. The KV cache grows linearly with the conversation, so a model that loads comfortably can still run out of memory halfway through a long document. On 96 GB, the context you get is whatever the weights leave behind — which is why two cards that run the same model can feel completely different in use. These are the largest contexts this card holds, at Q4_K_M weights.

ModelMax context, FP16 KVWith Q8 KV cache
Qwen3.8 27B 131,072 tokens 131,072 tokens
Llama 3.3 70B 131,072 tokens 131,072 tokens
Qwen3 32B 40,960 tokens · model limit 40,960 tokens · model limit
Gemma 4 31B 131,072 tokens 131,072 tokens

Every model in this table already reaches its limit at FP16, so quantising the KV cache buys nothing here. Capped at 128k tokens, or at the model's own native window where that is smaller, marked "model limit": free memory beyond that point buys nothing. Rows that hold far more context than their size suggests are hybrid, sliding-window or latent-attention models, which cache only a fraction of what a standard transformer does; each model page shows the working.

Where this card stops

Nothing in this list is out of reach. Every model tracked here fits at Q4_K_M with 8k of context, so the ceiling you will meet is context length and batch size, not parameter count.

The honest take

The M4 Max's ceiling of 128 GB with faster memory: 614 GB/s against 546 GB/s. That holds a 70B model at Q4 with a long context, on battery, at a speed that is fine for chat. Prompt processing is slower than on Nvidia, as on every Apple machine, so long documents pause before the first token. The cheaper M5 Max with the 32-core GPU is a different machine for this purpose: 36 GB and less bandwidth.

What it is good at

  • Up to 128 GB of unified memory in a laptop, with about 96 GB for the GPU.
  • 614 GB/s, about an eighth faster than the M4 Max.
  • MLX, llama.cpp, Ollama and LM Studio all run on it.

What it is not

  • Prompt processing lags Nvidia badly, which is felt most on long documents and agent loops.
  • No CUDA: much of the image, video and fine-tuning ecosystem is unavailable or slow.
  • Memory is fixed at purchase.

The thing people get wrong: Only the 40-core GPU version offers 48, 64 or 128 GB and the full 614 GB/s. The 32-core GPU M5 Max comes with 36 GB only.

Buy or rent

If you were buying a MacBook Pro anyway, the 128 GB configuration is the most capable portable machine for local models. If you were not, it is an expensive way to buy GPU memory, and it does not run CUDA.

What matters is the memory upgrade rather than the machine, and Apple prices that differently in every configuration. Put the figure you are actually looking at, with your electricity rate and your real monthly hours, into the build vs rent calculator. It returns the month owning becomes cheaper, or tells you it never does. Hours per month decides it far more often than the price does.

Nodegrove is the rented side of that comparison: a workspace that stays saved, with a GPU attached only while you are using it. It is the right answer when your hours are low or your needs change; buying is the right answer when they are high and stable.

Compared with the alternatives

  • MacBook Pro M4 Max — 96 GB · 546 GB/s · best fit Mistral Small 4 119B (MoE) at ~101 tok/s. A laptop that runs models a desktop GPU cannot hold.
  • Mac Studio M5 Ultra — 384 GB · 1,200 GB/s · best fit Qwen3.8 27B at ~52 tok/s. The most memory you can put on a desk: up to 512 GB at 1.2 TB/s, without CUDA.
  • DGX Spark — 126 GB · 273 GB/s · best fit Mistral Small 4 119B (MoE) at ~51 tok/s. NVIDIA's small Linux box with 128 GB the GPU can share: large models fit, but they are read at about a quarter of an RTX 4090's speed.

The unified-memory machines are compared head to head, on the same models, in DGX Spark vs Strix Halo vs Mac Studio.

All 20 cards are compared side by side on the GPU index, and the same numbers from the model's point of view are on each model fit table. To check one specific pairing rather than read a table, use the can-I-run-it checker.

Questions

Can a MacBook Pro M5 Max run a 70B model?

Yes. Llama 3.3 70B at Q4_K_M with 8k of context needs about 45.8 GB, inside the 96 GB this card can address, and it generates an estimated 10 tokens per second — slow; fine for batch jobs.

What is the best model to run on a MacBook Pro M5 Max?

Qwen3.8 27B. At Q4_K_M and 8k context it needs about 18 GB of the 96 GB available and generates an estimated 27 tokens per second, which is comfortable for chat. If quality matters more than size, Llama 3.3 70B fits at Q8 in about 81 GB.

How many tokens per second does a MacBook Pro M5 Max generate?

It depends entirely on the size of the model, because generating a token means reading every active weight out of memory. At 614 GB/s and 70% efficiency, this card produces an estimated 93 tokens per second on an 8B model at Q4, and about 114 on the largest model it holds, Mistral Small 4 119B (MoE). Reading speed is roughly 5 to 8 words per second, so anything above 25 feels immediate. Treat 93 as a ceiling on 614 GB/s rather than a measurement; the higher a figure, the further real runtimes fall below it.

Is 96 GB enough for running LLMs locally?

It runs 29 of the 29 models tracked here at Q4_K_M with 8k of context, up to 119B parameters. The honest test is not the model list but the context: Qwen3.8 27B on this card holds about 131,072 tokens before memory runs out.

MacBook Pro M5 Max or MacBook Pro M4 Max for local models?

Both address about the same memory, so they run the same models. On speed, this card is faster: 614 GB/s against 546 GB/s, and bandwidth is what sets chat speed.

M5 Max or M4 Max for local LLMs?

The same 128 GB ceiling, so the same models fit. The M5 Max reads memory at 614 GB/s against 546 GB/s, so it generates text about an eighth faster. If you already own a 128 GB M4 Max, that is not a reason to upgrade.

Which M5 Max should I buy for AI?

The 40-core GPU version. It is the only one offered with 48, 64 or 128 GB and the full 614 GB/s; the 32-core GPU version comes with 36 GB only. The GPU gets about 75% of whatever you buy, and memory cannot be upgraded later.

  • Memory: weights (parameters × bytes per parameter) + KV cache (2 × layers × KV heads × head dim × context × bytes for a standard transformer; sliding-window, hybrid and latent-attention models counted as they actually cache) + 0.5 GB runtime + 4% of weights. Full derivation in the VRAM calculator.
  • Speed: 70% × bandwidth ÷ active weight bytes, single stream, no batching, prompt processing excluded. A ceiling, not a measurement: the 11 of 29 figures marked ≤ are where real runtimes fall furthest below the number, because so few weights are read per token. See the tokens-per-second estimator.
  • Card specification: 128 GB, 614 GB/s — manufacturer figures; see Apple's own page for this machine. Apple unified memory: Apple publishes no GPU share; macOS's Metal limit is observed at about 75% of RAM on large-memory Macs, and every figure here uses that share.
  • Model architecture values come from each model's published config; the self-hosted LLM guide explains what each one changes.