What LLMs can a Mac Studio M3 Ultra run?

Which open models run on a Mac Studio M3 Ultra: what fits at Q4 and Q8, estimated tokens per second, how much context fits, and where the card stops.

Updated · estimates are labelled as estimates
On this page

512 GB unified 819 GB/s 2025 Apple unified memory

Mac Studio M3 Ultra (512 GB) addresses 384 GB — 75% of its 512 GB of unified memory, the share macOS is observed to hand the GPU by default, at 819 GB/s. Of the 29 open models tracked on this site, it runs 29 at Q4_K_M with 8k of context. The largest is Mistral Small 4 119B (MoE), needing about 72.5 GB and generating an estimated 152 tokens per second — faster than you can read.

2025's 512 GB Mac Studio: still the way to hold very large models, now mostly second-hand.

What a Mac Studio M3 Ultra runs, model by model

Every model at Q4_K_M with 8k of context, the setting most people actually use. This card has 364.8 GB to spend once the 5% safety margin comes off its 384 GB, and it reads that memory at 819 GB/s. Those two numbers decide the last two columns: the first says what loads, the second says how fast it answers. Both are estimates from stated formulas, and the VRAM calculator shows the working.

ModelSizeNeedsEst. tokens/sOn this card
Llama 3.2 3B 3.2B 3.4 GB ≤ 309 Fits
Qwen3.5 4B 4.7B 3.7 GB ≤ 210 Fits
Llama 3.1 8B 8B 6.4 GB ≤ 124 Fits
Qwen3 8B 8.2B 6.7 GB ≤ 121 Fits
Qwen3.5 9B 9.7B 6.7 GB ≤ 102 Fits
Gemma 4 12B 12B 8.2 GB 82 Fits
Gemma 3 12B 12.2B 8.7 GB 81 Fits
Ministral 3 14B 14B 10.3 GB 71 Fits
Qwen3 14B 14.8B 10.8 GB 67 Fits
Phi-4 14B 14.7B 11 GB 67 Fits
gpt-oss 20B 21B · 3.6B active 13.4 GB ≤ 275 Fits
Devstral Small 2 24B 24B 16.3 GB 41 Fits
Mistral Small 3.1 24B 24B 16.3 GB 41 Fits
Gemma 4 26B-A4B (MoE) 25.8B · 3.8B active 16.4 GB ≤ 260 Fits
Gemma 3 27B 27.4B 18.1 GB 36 Fits
Qwen3.8 27B 27.8B 18 GB 36 Fits
Qwen3 30B-A3B (MoE) 30.5B · 3.3B active 19.7 GB ≤ 300 Fits
GLM-4.7 Flash 30B-A3B (MoE) 31.2B · 3B active 19.8 GB ≤ 329 Fits
Gemma 4 31B 31.3B 20.9 GB 32 Fits
Nemotron 3.5 Lightning 30B-A3B (MoE) 31.6B · 3B active 19.7 GB ≤ 329 Fits
Qwen3 32B 32.8B 22.4 GB 30 Fits
Qwen2.5 Coder 32B 32.8B 22.4 GB 30 Fits
DeepSeek-R1 Distill Qwen 32B 32.8B 22.4 GB 30 Fits
Qwen3.6 35B-A3B (MoE) 36B · 3B active 22.4 GB ≤ 329 Fits
Llama 3.3 70B 70.6B 45.8 GB 14 Fits
DeepSeek-R1 Distill Llama 70B 70.6B 45.8 GB 14 Fits
Qwen3-Coder-Next 80B-A3B (MoE) 79.7B · 3B active 48.9 GB ≤ 329 Fits
gpt-oss 120B 117B · 5.1B active 71.4 GB ≤ 194 Fits
Mistral Small 4 119B (MoE) 119B · 6.5B active 72.5 GB ≤ 152 Fits

On this card that means anything at or under 364.8 GB counts as fitting, and anything above 326.4 GB is marked tight: it loads, but there is little left for the conversation, and a desktop on the same 512 GB can push it over.

One caveat specific to Apple silicon: these tokens-per-second figures describe generation, which unified memory handles well. Prompt processing — reading what you send before the first token comes back — is markedly slower than on an equivalent Nvidia card, and that is what you feel when you paste in a long document. Short prompts feel fast on this machine; 30,000-token ones do not.

The model to actually run on it

Qwen3.8 27B is the best use of this card: 18 GB of the 384 GB available, an estimated 36 tokens per second, comfortable for chat, and 131,072 tokens of context still available.

How this is chosen, since no benchmark is quoted: models are compared by size, a mixture-of-experts model counting at the geometric mean of its total and active parameters (a rule of thumb, not a measurement). The pick is the biggest class that fits with memory to spare and answers at 25 tokens per second or better, and within that class the most recent general-purpose release; coding and reasoning specialists are listed in the table but not recommended by default.

Mistral Small 4 119B (MoE) is the largest model the card will hold, at 72.5 GB, with room for 131,072 tokens of context. It is a mixture-of-experts model reading only 6.5B parameters per token, so it is the faster of the two at an estimated 152 tokens per second; what it gives up is the depth of a dense model that reads all of its weights for every token.

If quality matters more than parameter count, Llama 3.3 70B fits at Q8_0 in about 81 GB. Q8 is near-lossless; a smaller model at Q8 often beats a larger one squeezed into Q3.

How much context actually fits

Model size is the question people ask; context is the one that bites. The KV cache grows linearly with the conversation, so a model that loads comfortably can still run out of memory halfway through a long document. On 384 GB, the context you get is whatever the weights leave behind — which is why two cards that run the same model can feel completely different in use. These are the largest contexts this card holds, at Q4_K_M weights.

ModelMax context, FP16 KVWith Q8 KV cache
Qwen3.8 27B 131,072 tokens 131,072 tokens
Llama 3.3 70B 131,072 tokens 131,072 tokens
Qwen3 32B 40,960 tokens · model limit 40,960 tokens · model limit
Gemma 4 31B 131,072 tokens 131,072 tokens

Every model in this table already reaches its limit at FP16, so quantising the KV cache buys nothing here. Capped at 128k tokens, or at the model's own native window where that is smaller, marked "model limit": free memory beyond that point buys nothing. Rows that hold far more context than their size suggests are hybrid, sliding-window or latent-attention models, which cache only a fraction of what a standard transformer does; each model page shows the working.

Where this card stops

Nothing in this list is out of reach. Every model tracked here fits at Q4_K_M with 8k of context, so the ceiling you will meet is context length and batch size, not parameter count.

The honest take

It launched with up to 512 GB at 819 GB/s, the first desktop that could hold the biggest open mixture-of-experts models at all. The M5 Ultra replaced it in September 2026 and reads memory about half as fast again, but a 512 GB M3 Ultra holds exactly what a 512 GB M5 Ultra holds. Apple's current spec page lists it only up to 256 GB.

What it is good at

  • Up to 512 GB of unified memory, about 384 GB of it for the GPU.
  • 819 GB/s, a little under an RTX 3090's 936 GB/s, for generation.
  • Silent and low-power for what it holds.

What it is not

  • Prompt processing is far slower than on Nvidia.
  • No CUDA.
  • Slower than the M5 Ultra that replaced it, at the same capacity.

The thing people get wrong: Apple's spec page for this Mac Studio now lists 96 GB configurable to 256 GB; the 512 GB option, offered at launch, has gone from it. A 512 GB unit is an original configuration with the 32-core CPU and 80-core GPU, so check that before buying one second-hand.

Buy or rent

Worth it second-hand at 512 GB if very large models are the point and you do not need CUDA. At 96 or 256 GB, compare it with a current M5 Max or M5 Ultra before deciding.

What matters is the memory upgrade rather than the machine, and Apple prices that differently in every configuration. Put the figure you are actually looking at, with your electricity rate and your real monthly hours, into the build vs rent calculator. It returns the month owning becomes cheaper, or tells you it never does. Hours per month decides it far more often than the price does.

Nodegrove is the rented side of that comparison: a workspace that stays saved, with a GPU attached only while you are using it. It is the right answer when your hours are low or your needs change; buying is the right answer when they are high and stable.

Compared with the alternatives

  • Mac Studio M5 Ultra — 384 GB · 1,200 GB/s · best fit Qwen3.8 27B at ~52 tok/s. The most memory you can put on a desk: up to 512 GB at 1.2 TB/s, without CUDA.
  • Mac Studio M2 Ultra — 144 GB · 800 GB/s · best fit Qwen3.8 27B at ~35 tok/s. The quiet way to hold a very large model, as long as you are patient with long prompts.
  • DGX Spark — 126 GB · 273 GB/s · best fit Mistral Small 4 119B (MoE) at ~51 tok/s. NVIDIA's small Linux box with 128 GB the GPU can share: large models fit, but they are read at about a quarter of an RTX 4090's speed.

The unified-memory machines are compared head to head, on the same models, in DGX Spark vs Strix Halo vs Mac Studio.

All 20 cards are compared side by side on the GPU index, and the same numbers from the model's point of view are on each model fit table. To check one specific pairing rather than read a table, use the can-I-run-it checker.

Questions

Can a Mac Studio M3 Ultra run a 70B model?

Yes. Llama 3.3 70B at Q4_K_M with 8k of context needs about 45.8 GB, inside the 384 GB this card can address, and it generates an estimated 14 tokens per second — usable, a little slow.

What is the best model to run on a Mac Studio M3 Ultra?

Qwen3.8 27B. At Q4_K_M and 8k context it needs about 18 GB of the 384 GB available and generates an estimated 36 tokens per second, which is comfortable for chat. If quality matters more than size, Llama 3.3 70B fits at Q8 in about 81 GB.

How many tokens per second does a Mac Studio M3 Ultra generate?

It depends entirely on the size of the model, because generating a token means reading every active weight out of memory. At 819 GB/s and 70% efficiency, this card produces an estimated 124 tokens per second on an 8B model at Q4, and about 152 on the largest model it holds, Mistral Small 4 119B (MoE). Reading speed is roughly 5 to 8 words per second, so anything above 25 feels immediate. Treat 124 as a ceiling on 819 GB/s rather than a measurement; the higher a figure, the further real runtimes fall below it.

Is 384 GB enough for running LLMs locally?

It runs 29 of the 29 models tracked here at Q4_K_M with 8k of context, up to 119B parameters. The honest test is not the model list but the context: Qwen3.8 27B on this card holds about 131,072 tokens before memory runs out.

Mac Studio M3 Ultra or Mac Studio M5 Ultra for local models?

Both address about the same memory, so they run the same models. On speed, the Mac Studio M5 Ultra is faster: 1,200 GB/s against 819 GB/s, and bandwidth is what sets chat speed.

Is the M3 Ultra still worth it for local LLMs?

For capacity, yes: a 512 GB unit holds what a 512 GB M5 Ultra holds. It generates more slowly, at 819 GB/s against 1,200 GB/s, and Apple's current spec page no longer lists the 512 GB configuration.

How much of an M3 Ultra's memory can the GPU use?

About 75% by default: roughly 384 GB of 512, or 192 GB of 256. Apple does not publish a figure; that is the macOS limit observed on large-memory Macs.

  • Memory: weights (parameters × bytes per parameter) + KV cache (2 × layers × KV heads × head dim × context × bytes for a standard transformer; sliding-window, hybrid and latent-attention models counted as they actually cache) + 0.5 GB runtime + 4% of weights. Full derivation in the VRAM calculator.
  • Speed: 70% × bandwidth ÷ active weight bytes, single stream, no batching, prompt processing excluded. A ceiling, not a measurement: the 14 of 29 figures marked ≤ are where real runtimes fall furthest below the number, because so few weights are read per token. See the tokens-per-second estimator.
  • Card specification: 512 GB, 819 GB/s — manufacturer figures; see Apple's own page for this machine. The bandwidth figure: Apple's M3 Ultra Mac Studio tech specs. Apple unified memory: Apple publishes no GPU share; macOS's Metal limit is observed at about 75% of RAM on large-memory Macs, and every figure here uses that share.
  • Model architecture values come from each model's published config; the self-hosted LLM guide explains what each one changes.