128 GB unified 256 GB/s 2025 Sold new
Ryzen AI Max+ 395 (128 GB) addresses 96 GB of its 128 GB of unified memory, worked out from what AMD documents (see the sources below), at 256 GB/s. Of the 29 open models tracked on this site, it runs 29 at Q4_K_M with 8k of context. The largest is Mistral Small 4 119B (MoE), needing about 72.5 GB and generating an estimated 48 tokens per second — comfortable for chat.
AMD's 128 GB laptop and mini-PC chip: up to 96 GB for the GPU on Windows, and an x86 route to large mixture-of-experts models.
What a Ryzen AI Max+ 395 runs, model by model
Every model at Q4_K_M with 8k of context, the setting most people actually use. This card has 91.2 GB to spend once the 5% safety margin comes off its 96 GB, and it reads that memory at 256 GB/s. Those two numbers decide the last two columns: the first says what loads, the second says how fast it answers. Both are estimates from stated formulas, and the VRAM calculator shows the working.
| Model | Size | Needs | Est. tokens/s | On this card |
|---|---|---|---|---|
| Llama 3.2 3B | 3.2B | 3.4 GB | 97 | Fits |
| Qwen3.5 4B | 4.7B | 3.7 GB | 66 | Fits |
| Llama 3.1 8B | 8B | 6.4 GB | 39 | Fits |
| Qwen3 8B | 8.2B | 6.7 GB | 38 | Fits |
| Qwen3.5 9B | 9.7B | 6.7 GB | 32 | Fits |
| Gemma 4 12B | 12B | 8.2 GB | 26 | Fits |
| Gemma 3 12B | 12.2B | 8.7 GB | 25 | Fits |
| Ministral 3 14B | 14B | 10.3 GB | 22 | Fits |
| Qwen3 14B | 14.8B | 10.8 GB | 21 | Fits |
| Phi-4 14B | 14.7B | 11 GB | 21 | Fits |
| gpt-oss 20B | 21B · 3.6B active | 13.4 GB | 86 | Fits |
| Devstral Small 2 24B | 24B | 16.3 GB | 13 | Fits |
| Mistral Small 3.1 24B | 24B | 16.3 GB | 13 | Fits |
| Gemma 4 26B-A4B (MoE) | 25.8B · 3.8B active | 16.4 GB | 81 | Fits |
| Gemma 3 27B | 27.4B | 18.1 GB | 11 | Fits |
| Qwen3.8 27B | 27.8B | 18 GB | 11 | Fits |
| Qwen3 30B-A3B (MoE) | 30.5B · 3.3B active | 19.7 GB | 94 | Fits |
| GLM-4.7 Flash 30B-A3B (MoE) | 31.2B · 3B active | 19.8 GB | ≤ 103 | Fits |
| Gemma 4 31B | 31.3B | 20.9 GB | 10 | Fits |
| Nemotron 3.5 Lightning 30B-A3B (MoE) | 31.6B · 3B active | 19.7 GB | ≤ 103 | Fits |
| Qwen3 32B | 32.8B | 22.4 GB | 9 | Fits |
| Qwen2.5 Coder 32B | 32.8B | 22.4 GB | 9 | Fits |
| DeepSeek-R1 Distill Qwen 32B | 32.8B | 22.4 GB | 9 | Fits |
| Qwen3.6 35B-A3B (MoE) | 36B · 3B active | 22.4 GB | ≤ 103 | Fits |
| Llama 3.3 70B | 70.6B | 45.8 GB | 4 | Fits |
| DeepSeek-R1 Distill Llama 70B | 70.6B | 45.8 GB | 4 | Fits |
| Qwen3-Coder-Next 80B-A3B (MoE) | 79.7B · 3B active | 48.9 GB | ≤ 103 | Fits |
| gpt-oss 120B | 117B · 5.1B active | 71.4 GB | 61 | Fits |
| Mistral Small 4 119B (MoE) | 119B · 6.5B active | 72.5 GB | 48 | Fits |
On this card that means anything at or under 91.2 GB counts as fitting, and anything above 81.6 GB is marked tight: it loads, but there is little left for the conversation, and a desktop on the same 128 GB can push it over.
The model to actually run on it
Mistral Small 4 119B (MoE) is the best use of this card: 72.5 GB of the 96 GB available, an estimated 48 tokens per second, comfortable for chat, and 131,072 tokens of context still available. It is a mixture-of-experts model, so all 119B parameters sit in memory but only 6.5B are read per token — which is why it is quick for its size.
How this is chosen, since no benchmark is quoted: models are compared by size, a mixture-of-experts model counting at the geometric mean of its total and active parameters (a rule of thumb, not a measurement). The pick is the biggest class that fits with memory to spare and answers at 25 tokens per second or better, and within that class the most recent general-purpose release; coding and reasoning specialists are listed in the table but not recommended by default.
If quality matters more than parameter count, Llama 3.3 70B fits at Q8_0 in about 81 GB. Q8 is near-lossless; a smaller model at Q8 often beats a larger one squeezed into Q3.
How much context actually fits
Model size is the question people ask; context is the one that bites. The KV cache grows linearly with the conversation, so a model that loads comfortably can still run out of memory halfway through a long document. On 96 GB, the context you get is whatever the weights leave behind — which is why two cards that run the same model can feel completely different in use. These are the largest contexts this card holds, at Q4_K_M weights.
| Model | Max context, FP16 KV | With Q8 KV cache |
|---|---|---|
| Mistral Small 4 119B (MoE) | 131,072 tokens | 131,072 tokens |
| Llama 3.3 70B | 131,072 tokens | 131,072 tokens |
| Qwen3 32B | 40,960 tokens · model limit | 40,960 tokens · model limit |
| Gemma 4 31B | 131,072 tokens | 131,072 tokens |
Every model in this table already reaches its limit at FP16, so quantising the KV cache buys nothing here. Capped at 128k tokens, or at the model's own native window where that is smaller, marked "model limit": free memory beyond that point buys nothing. Rows that hold far more context than their size suggests are hybrid, sliding-window or latent-attention models, which cache only a fraction of what a standard transformer does; each model page shows the working.
Where this card stops
Nothing in this list is out of reach. Every model tracked here fits at Q4_K_M with 8k of context, so the ceiling you will meet is context length and batch size, not parameter count.
The honest take
The same bargain as the DGX Spark, made by AMD for ordinary PCs: lots of memory, modest speed. Up to 96 GB can go to the GPU on Windows, and more on Linux, which holds models far beyond any 24 GB card. At 256 GB/s it reads them slowly, so mixture-of-experts models suit it and dense 70B models do not. It runs Windows and Linux on x86, through llama.cpp, LM Studio and ROCm rather than CUDA.
What it is good at
- Up to 96 GB of graphics memory on Windows, set with AMD Variable Graphics Memory.
- An x86 machine that runs Windows or Linux, in laptops and mini-PCs from several makers.
- LM Studio and llama.cpp support it directly, and AMD documents ROCm on Linux for this chip.
What it is not
- 256 GB/s: every token of a dense model reads all its weights at that speed.
- No CUDA. Fine-tuning and much image and video tooling assume it.
- On Linux the GPU sees about half the RAM until the shared memory limit is raised, and ROCm needs a recent kernel.
The thing people get wrong: The 128 GB is not all GPU memory. On Windows up to 96 GB can be made graphics memory, and whatever you assign is taken away from the CPU. The power limit is chosen by the laptop or mini-PC maker, anywhere from 45 to 120 W, so two machines with this chip can perform quite differently.
Buy or rent
It makes sense if you want one x86 machine, laptop or desktop, that also holds large mixture-of-experts models. For dense models above 30B it is slow: a used 24 GB card generates several times faster on anything that fits in 24 GB. Check the power limit of the specific machine before you buy.
Retail prices drift, and a figure written today would be wrong within a month, so this page does not carry one. Put the figure you are actually looking at, with your electricity rate and your real monthly hours, into the build vs rent calculator. It returns the month owning becomes cheaper, or tells you it never does. Hours per month decides it far more often than the price does.
Nodegrove is the rented side of that comparison: a workspace that stays saved, with a GPU attached only while you are using it. It is the right answer when your hours are low or your needs change; buying is the right answer when they are high and stable.
Compared with the alternatives
- DGX Spark — 126 GB · 273 GB/s · best fit Mistral Small 4 119B (MoE) at ~51 tok/s. NVIDIA's small Linux box with 128 GB the GPU can share: large models fit, but they are read at about a quarter of an RTX 4090's speed.
- MacBook Pro M5 Max — 96 GB · 614 GB/s · best fit Qwen3.8 27B at ~27 tok/s. The laptop for large models in 2026: 128 GB at 614 GB/s, if you buy the 40-core GPU.
- RTX 3090 — 24 GB · 936 GB/s · best fit Qwen3.8 27B at ~41 tok/s. The value benchmark for local models, and it has been for years.
The unified-memory machines are compared head to head, on the same models, in DGX Spark vs Strix Halo vs Mac Studio.
All 20 cards are compared side by side on the GPU index, and the same numbers from the model's point of view are on each model fit table. To check one specific pairing rather than read a table, use the can-I-run-it checker.
Questions
Can a Ryzen AI Max+ 395 run a 70B model?
Yes. Llama 3.3 70B at Q4_K_M with 8k of context needs about 45.8 GB, inside the 96 GB this card can address, and it generates an estimated 4 tokens per second — too slow for interactive use.
What is the best model to run on a Ryzen AI Max+ 395?
Mistral Small 4 119B (MoE). At Q4_K_M and 8k context it needs about 72.5 GB of the 96 GB available and generates an estimated 48 tokens per second, which is comfortable for chat. If quality matters more than size, Llama 3.3 70B fits at Q8 in about 81 GB.
How many tokens per second does a Ryzen AI Max+ 395 generate?
It depends entirely on the size of the model, because generating a token means reading every active weight out of memory. At 256 GB/s and 70% efficiency, this card produces an estimated 39 tokens per second on an 8B model at Q4, and about 48 on the largest model it holds, Mistral Small 4 119B (MoE). Reading speed is roughly 5 to 8 words per second, so anything above 25 feels immediate. Treat 39 as a ceiling on 256 GB/s rather than a measurement; the higher a figure, the further real runtimes fall below it.
Is 96 GB enough for running LLMs locally?
It runs 29 of the 29 models tracked here at Q4_K_M with 8k of context, up to 119B parameters. The honest test is not the model list but the context: Mistral Small 4 119B (MoE) on this card holds about 131,072 tokens before memory runs out.
Ryzen AI Max+ 395 or DGX Spark for local models?
Both address about the same memory, so they run the same models. On speed, the DGX Spark is faster: 273 GB/s against 256 GB/s, and bandwidth is what sets chat speed.
How much memory can the Ryzen AI Max+ 395's GPU use?
On Windows, up to 96 GB of a 128 GB machine can be set as graphics memory with AMD Variable Graphics Memory, and AMD says workloads perform best inside that. On Linux, AMD recommends a small BIOS reservation and raising the shared memory limit instead: the default is about half of RAM, and AMD's own guide raises it to 120 GB. Every figure on this page uses 96 GB.
Is Strix Halo good for local LLMs?
For mixture-of-experts models, yes: they hold tens of billions of parameters but read only a few billion per token, so capacity matters more than its 256 GB/s. For dense 70B models it is slow, because every token reads all of the weights at that speed.
- Memory: weights (parameters × bytes per parameter) + KV cache (2 × layers × KV heads × head dim × context × bytes for a standard transformer; sliding-window, hybrid and latent-attention models counted as they actually cache) + 0.5 GB runtime + 4% of weights. Full derivation in the VRAM calculator.
- Speed: 70% × bandwidth ÷ active weight bytes, single stream, no batching, prompt processing excluded. A ceiling, not a measurement: the 4 of 29 figures marked ≤ are where real runtimes fall furthest below the number, because so few weights are read per token. See the tokens-per-second estimator.
- Card specification: 128 GB, 256 GB/s — manufacturer figures; see AMD's own page for this machine. The bandwidth figure: AMD's Ryzen AI Halo page for the same chip. Unified memory: AMD lets up to 96 GB of the 128 GB be set aside as graphics memory (Variable Graphics Memory), and says workloads run best inside it. Figures here use 96 GB. On Linux AMD documents raising the shared limit further, to 120 GB in its own guide. Source: AMD's Variable Graphics Memory FAQ.
- Model architecture values come from each model's published config; the self-hosted LLM guide explains what each one changes.