What LLMs can an RX 7900 XTX run?

Which open models run on an RX 7900 XTX: what fits at Q4 and Q8, estimated tokens per second, how much context fits, and where the card stops.

Updated · estimates are labelled as estimates
On this page

24 GB VRAM 960 GB/s 355 W 2022 Sold new

RX 7900 XTX (24 GB) addresses 24 GB at 960 GB/s. Of the 29 open models tracked on this site, it runs 24 at Q4_K_M with 8k of context. The largest is Qwen3.6 35B-A3B (MoE), needing about 22.4 GB and generating an estimated 386 tokens per second — faster than you can read.

AMD's 24 GB card: 3090-class memory and bandwidth, without CUDA.

What an RX 7900 XTX runs, model by model

Every model at Q4_K_M with 8k of context, the setting most people actually use. This card has 22.8 GB to spend once the 5% safety margin comes off its 24 GB, and it reads that memory at 960 GB/s. Those two numbers decide the last two columns: the first says what loads, the second says how fast it answers. Both are estimates from stated formulas, and the VRAM calculator shows the working.

ModelSizeNeedsEst. tokens/sOn this card
Llama 3.2 3B 3.2B 3.4 GB ≤ 362 Fits
Qwen3.5 4B 4.7B 3.7 GB ≤ 247 Fits
Llama 3.1 8B 8B 6.4 GB ≤ 145 Fits
Qwen3 8B 8.2B 6.7 GB ≤ 141 Fits
Qwen3.5 9B 9.7B 6.7 GB ≤ 119 Fits
Gemma 4 12B 12B 8.2 GB 97 Fits
Gemma 3 12B 12.2B 8.7 GB 95 Fits
Ministral 3 14B 14B 10.3 GB 83 Fits
Qwen3 14B 14.8B 10.8 GB 78 Fits
Phi-4 14B 14.7B 11 GB 79 Fits
gpt-oss 20B 21B · 3.6B active 13.4 GB ≤ 322 Fits
Devstral Small 2 24B 24B 16.3 GB 48 Fits
Mistral Small 3.1 24B 24B 16.3 GB 48 Fits
Gemma 4 26B-A4B (MoE) 25.8B · 3.8B active 16.4 GB ≤ 305 Fits
Gemma 3 27B 27.4B 18.1 GB 42 Fits
Qwen3.8 27B 27.8B 18 GB 42 Fits
Qwen3 30B-A3B (MoE) 30.5B · 3.3B active 19.7 GB ≤ 351 Fits
GLM-4.7 Flash 30B-A3B (MoE) 31.2B · 3B active 19.8 GB ≤ 386 Fits
Gemma 4 31B 31.3B 20.9 GB 37 Tight fit
Nemotron 3.5 Lightning 30B-A3B (MoE) 31.6B · 3B active 19.7 GB ≤ 386 Fits
Qwen3 32B 32.8B 22.4 GB 35 Tight fit
Qwen2.5 Coder 32B 32.8B 22.4 GB 35 Tight fit
DeepSeek-R1 Distill Qwen 32B 32.8B 22.4 GB 35 Tight fit
Qwen3.6 35B-A3B (MoE) 36B · 3B active 22.4 GB ≤ 386 Tight fit
Llama 3.3 70B 70.6B 45.8 GB – No
DeepSeek-R1 Distill Llama 70B 70.6B 45.8 GB – No
Qwen3-Coder-Next 80B-A3B (MoE) 79.7B · 3B active 48.9 GB – No
gpt-oss 120B 117B · 5.1B active 71.4 GB – No
Mistral Small 4 119B (MoE) 119B · 6.5B active 72.5 GB – No

On this card that means anything at or under 22.8 GB counts as fitting, and anything above 20.4 GB is marked tight: it loads, but there is little left for the conversation, and a desktop on the same 24 GB can push it over.

The model to actually run on it

Qwen3.8 27B is the best use of this card: 18 GB of the 24 GB available, an estimated 42 tokens per second, comfortable for chat, and 81,920 tokens of context still available.

How this is chosen, since no benchmark is quoted: models are compared by size, a mixture-of-experts model counting at the geometric mean of its total and active parameters (a rule of thumb, not a measurement). The pick is the biggest class that fits with memory to spare and answers at 25 tokens per second or better, and within that class the most recent general-purpose release; coding and reasoning specialists are listed in the table but not recommended by default.

Qwen3.6 35B-A3B (MoE) is the largest model the card will hold, at 22.4 GB, but it is the wrong daily driver: filling 93% of memory leaves room for only 25,600 tokens of conversation. More parameters are not worth a context window that runs out mid-document. It is a mixture-of-experts model reading only 3B parameters per token, so it is the faster of the two at an estimated 386 tokens per second; what it gives up is the depth of a dense model that reads all of its weights for every token.

If quality matters more than parameter count, Gemma 4 12B fits at Q8_0 in about 14.2 GB. Q8 is near-lossless; a smaller model at Q8 often beats a larger one squeezed into Q3.

How much context actually fits

Model size is the question people ask; context is the one that bites. The KV cache grows linearly with the conversation, so a model that loads comfortably can still run out of memory halfway through a long document. On 24 GB, the context you get is whatever the weights leave behind — which is why two cards that run the same model can feel completely different in use. These are the largest contexts this card holds, at Q4_K_M weights.

ModelMax context, FP16 KVWith Q8 KV cache
Qwen3.8 27B 81,920 tokens 131,072 tokens
Qwen3 32B 9,216 tokens 18,432 tokens
Gemma 4 31B 30,720 tokens 72,704 tokens
Gemma 3 27B 64,512 tokens 131,072 tokens

Quantising the KV cache to Q8 roughly doubles what you can hold — Qwen3.8 27B on this card goes from 81,920 to 131,072 tokens — at a quality cost most people never notice. Capped at 128k tokens, or at the model's own native window where that is smaller, marked "model limit": free memory beyond that point buys nothing. Rows that hold far more context than their size suggests are hybrid, sliding-window or latent-attention models, which cache only a fraction of what a standard transformer does; each model page shows the working.

Where this card stops

The first model out of reach is Llama 3.3 70B: about 45.8 GB at Q4_K_M and 8k context, against 24 GB of usable memory. Dropping to Q3_K_M would need about 37.7 GB, which still does not fit. Offloading the remainder to system RAM works and is ten to fifty times slower; it is a way to see a model run, not a way to use one.

The honest take

On paper it is a 3090: 24 GB at 960 GB/s, so the same models fit and the bandwidth formula gives them about the same speed. The difference is software. It runs language models well through llama.cpp, LM Studio and Ollama, using ROCm or Vulkan, but fine-tuning and most image and video tooling are written for CUDA first.

What it is good at

  • 24 GB at 960 GB/s: the same tier as an RTX 3090 or 4090 for what fits.
  • llama.cpp, LM Studio and Ollama all run on it, through ROCm or Vulkan.
  • AMD's ROCm 10 compatibility list now includes it on Windows 11 as well as Linux.

What it is not

  • No CUDA. Tools that only ship CUDA builds, which is most fine-tuning and much of the image and video ecosystem, need a ROCm port or do not run.
  • No FP8 support on this generation, so FP8 checkpoints are converted rather than run natively.
  • 355 W of board power and a large cooler.

The thing people get wrong: AMD's page prints two bandwidth figures. The much larger "effective" one counts the on-chip cache, which model weights do not fit in; for language models, 960 GB/s is the number that matters.

Buy or rent

Worth it if what you run is llama.cpp-based chat and you are comfortable checking each tool for ROCm support. If you want everything to work first time, including image generation and fine-tuning, the CUDA ecosystem decides it for a 3090 or 4090.

Retail prices drift, and a figure written today would be wrong within a month, so this page does not carry one. Put the figure you are actually looking at, with your electricity rate and your real monthly hours, into the build vs rent calculator — it already knows this card draws 355 W under load. It returns the month owning becomes cheaper, or tells you it never does. Hours per month decides it far more often than the price does.

Nodegrove is the rented side of that comparison: a workspace that stays saved, with a GPU attached only while you are using it. It is the right answer when your hours are low or your needs change; buying is the right answer when they are high and stable.

Compared with the alternatives

  • RTX 3090 — 24 GB · 936 GB/s · best fit Qwen3.8 27B at ~41 tok/s. The value benchmark for local models, and it has been for years.
  • RTX 4090 — 24 GB · 1,008 GB/s · best fit Qwen3.8 27B at ~44 tok/s. The card most local-model advice is implicitly written for.
  • RTX 5090 — 32 GB · 1,792 GB/s · best fit Qwen3.8 27B at ~78 tok/s. The first consumer card whose memory bandwidth changes what is comfortable.

All 20 cards are compared side by side on the GPU index, and the same numbers from the model's point of view are on each model fit table. To check one specific pairing rather than read a table, use the can-I-run-it checker.

Questions

Can an RX 7900 XTX run a 70B model?

Not at Q4_K_M in one card. Llama 3.3 70B needs about 45.8 GB at 8k context and this card can address 24 GB. Your options are a smaller model, a harsher quantisation with very little context, splitting across two cards, or renting a larger one for the hours you need it.

What is the best model to run on an RX 7900 XTX?

Qwen3.8 27B. At Q4_K_M and 8k context it needs about 18 GB of the 24 GB available and generates an estimated 42 tokens per second, which is comfortable for chat. If quality matters more than size, Gemma 4 12B fits at Q8 in about 14.2 GB.

How many tokens per second does an RX 7900 XTX generate?

It depends entirely on the size of the model, because generating a token means reading every active weight out of memory. At 960 GB/s and 70% efficiency, this card produces an estimated 145 tokens per second on an 8B model at Q4, and about 386 on the largest model it holds, Qwen3.6 35B-A3B (MoE). Reading speed is roughly 5 to 8 words per second, so anything above 25 feels immediate. Treat 145 as a ceiling on 960 GB/s rather than a measurement; the higher a figure, the further real runtimes fall below it.

Is 24 GB enough for running LLMs locally?

It runs 24 of the 29 models tracked here at Q4_K_M with 8k of context, up to 36B parameters. The honest test is not the model list but the context: Qwen3.8 27B on this card holds about 81,920 tokens before memory runs out.

RX 7900 XTX or RTX 3090 for local models?

Both address about the same memory, so they run the same models. On speed, this card is faster: 960 GB/s against 936 GB/s, and bandwidth is what sets chat speed.

Is the RX 7900 XTX good for local LLMs?

For running models in llama.cpp, LM Studio or Ollama, yes: 24 GB at 960 GB/s puts it level with an RTX 3090 on what fits and, by the bandwidth formula, on how fast it generates. Fine-tuning and most image and video tools still assume CUDA, so check your tools before you buy.

RX 7900 XTX or RTX 3090?

The same 24 GB and nearly the same bandwidth, 960 GB/s against 936 GB/s, so the same models fit at about the same estimated speed. The 3090 runs CUDA, which most of the ecosystem is built on; the 7900 XTX relies on ROCm or Vulkan, which covers chat well and the rest unevenly.

  • Memory: weights (parameters × bytes per parameter) + KV cache (2 × layers × KV heads × head dim × context × bytes for a standard transformer; sliding-window, hybrid and latent-attention models counted as they actually cache) + 0.5 GB runtime + 4% of weights. Full derivation in the VRAM calculator.
  • Speed: 70% × bandwidth ÷ active weight bytes, single stream, no batching, prompt processing excluded. A ceiling, not a measurement: the 11 of 24 figures marked ≤ are where real runtimes fall furthest below the number, because so few weights are read per token. See the tokens-per-second estimator.
  • Card specification: 24 GB, 960 GB/s, 355 W — manufacturer figures; see AMD's own page for this card.
  • Model architecture values come from each model's published config; the self-hosted LLM guide explains what each one changes.