What LLMs can an RTX 5090 run?

Which open models run on an RTX 5090: what fits at Q4 and Q8, estimated tokens per second, how much context fits, and where the card stops.

Updated 9 September 2026 · estimates are labelled as estimates

32 GB VRAM 1,792 GB/s 575 W 2025 Sold new

RTX 5090 (32 GB) addresses 32 GB at 1,792 GB/s. Of the 16 open models tracked on this site, it runs 13 at Q4_K_M with 8k of context. The largest is Qwen3 32B, needing about 22.4 GB and generating an estimated 66 tokens per second — faster than you can read.

The first consumer card whose memory bandwidth changes what is comfortable.

What an RTX 5090 runs, model by model

Every model at Q4_K_M with 8k of context, the setting most people actually use. This card has 30.4 GB to spend once the 5% safety margin comes off its 32 GB, and it reads that memory at 1,792 GB/s. Those two numbers decide the last two columns: the first says what loads, the second says how fast it answers. Both are estimates from stated formulas, and the VRAM calculator shows the working.

ModelSizeNeedsEst. tokens/sOn this card
Llama 3.2 3B 3.2B 3.4 GB 676 fits
Llama 3.1 8B 8B 6.4 GB 270 fits
Qwen3 8B 8.2B 6.7 GB 264 fits
Gemma 3 12B 12.2B 11.1 GB 177 fits
Qwen3 14B 14.8B 10.8 GB 146 fits
Phi-4 14B 14.7B 11 GB 147 fits
gpt-oss 20B 21B · 3.6B active 13.6 GB 601 fits
Mistral Small 3.1 24B 24B 16.3 GB 90 fits
Gemma 3 27B 27.4B 21.2 GB 79 fits
Qwen3 30B-A3B (MoE) 30.5B · 3.3B active 19.7 GB 655 fits
Qwen3 32B 32.8B 22.4 GB 66 fits
Qwen2.5 Coder 32B 32.8B 22.4 GB 66 fits
DeepSeek-R1 Distill Qwen 32B 32.8B 22.4 GB 66 fits
Llama 3.3 70B 70.6B 45.8 GB no
DeepSeek-R1 Distill Llama 70B 70.6B 45.8 GB no
gpt-oss 120B 117B · 5.1B active 71.7 GB no

On this card that means anything at or under 30.4 GB counts as fitting, and anything above 27.2 GB is marked tight: it loads, but there is little left for the conversation, and a desktop on the same 32 GB can push it over.

The model to actually run on it

Qwen3 32B is the best use of this card: 22.4 GB of the 32 GB available, an estimated 66 tokens per second, faster than you can read, and 37,888 tokens of context still available.

If quality matters more than parameter count, Mistral Small 3.1 24B fits at Q8_0 in about 28.3 GB. Q8 is near-lossless; a smaller model at Q8 often beats a larger one squeezed into Q3.

How much context actually fits

Model size is the question people ask; context is the one that bites. The KV cache grows linearly with the conversation, so a model that loads comfortably can still run out of memory halfway through a long document. On 32 GB, the context you get is whatever the weights leave behind — which is why two cards that run the same model can feel completely different in use. These are the largest contexts this card holds, at Q4_K_M weights.

ModelMax context, FP16 KVWith Q8 KV cache
Qwen3 32B 37,888 tokens 76,800 tokens
Qwen2.5 Coder 32B 37,888 tokens 76,800 tokens
DeepSeek-R1 Distill Qwen 32B 37,888 tokens 76,800 tokens
Qwen3 30B-A3B (MoE) 116,736 tokens 131,072 tokens

Quantising the KV cache to Q8 roughly doubles what you can hold — Qwen3 32B on this card goes from 37,888 to 76,800 tokens — at a quality cost most people never notice. Capped at 128k tokens here; a model's own architecture may stop lower, and sliding-window models such as Gemma 3 use less KV than the formula assumes.

Where this card stops

The first model out of reach is Llama 3.3 70B: about 45.8 GB at Q4_K_M and 8k context, against 32 GB of usable memory. Dropping to Q3_K_M would need about 37.7 GB, which still does not fit. Offloading the remainder to system RAM works and is ten to fifty times slower; it is a way to see a model run, not a way to use one.

The honest take

Thirty-two gigabytes is a modest capacity step; 1792 GB/s is not. The 5090 generates text roughly 78% faster than a 4090 on the same model, which moves 32B-class models from "quick" to "instant" and makes larger models at aggressive quantisation genuinely usable. It is also 575 W, which is a real constraint in a room you sit in.

What it is good at

  • 1792 GB/s is the highest bandwidth outside the datacenter, and bandwidth is what sets chat speed.
  • 32 GB clears the 24 GB cliff: 32B models fit at Q4 with long context, and Q5 becomes an option.
  • GDDR7 and Blackwell compute make it the strongest consumer card for image and video generation as well.

What it is not

  • 575 W. In a small room this is a heater, and it needs a power supply chosen for it.
  • 32 GB still does not fit a 70B model at Q4 with useful context.
  • Newest generation means the occasional runtime or driver rough edge that older cards no longer have.

The thing people get wrong: The 8 GB of extra memory over a 4090 is not why this card is faster. Text generation speed is bandwidth divided by weight size, and the 5090 moves 78% more memory per second. That single number explains almost all of the difference.

Buy or rent

The case for buying is strongest here if you run models daily and value the time. For occasional use, the power draw and the outlay both argue for renting a bigger card only when you need it.

Retail prices drift, and a figure written today would be wrong within a month, so this page does not carry one. Put the figure you are actually looking at, with your electricity rate and your real monthly hours, into the build vs rent calculator — it already knows this card draws 575 W under load. It returns the month owning becomes cheaper, or tells you it never does. Hours per month decides it far more often than the price does.

Nodegrove is the rented side of that comparison: a workspace that stays saved, with a GPU attached only while you are using it. It is the right answer when your hours are low or your needs change; buying is the right answer when they are high and stable.

Compared with the alternatives

  • RTX 409024 GB · 1,008 GB/s · best fit Qwen3 30B-A3B (MoE) at ~369 tok/s. The card most local-model advice is implicitly written for.
  • RTX PRO 6000 Blackwell96 GB · 1,792 GB/s · best fit gpt-oss 120B at ~424 tok/s. Ninety-six gigabytes and 5090-class bandwidth on one card.
  • RTX 309024 GB · 936 GB/s · best fit Qwen3 30B-A3B (MoE) at ~342 tok/s. The value benchmark for local models, and it has been for years.

All 13 cards are compared side by side on the GPU index, and the same numbers from the model's point of view are on each model fit table. To check one specific pairing rather than read a table, use the can-I-run-it checker.

Questions

Can an RTX 5090 run a 70B model?

Not at Q4_K_M in one card. Llama 3.3 70B needs about 45.8 GB at 8k context and this card can address 32 GB. Your options are a smaller model, a harsher quantisation with very little context, splitting across two cards, or renting a larger one for the hours you need it.

What is the best model to run on an RTX 5090?

Qwen3 32B. At Q4_K_M and 8k context it needs about 22.4 GB of the 32 GB available and generates an estimated 66 tokens per second, which is faster than you can read. If quality matters more than size, Mistral Small 3.1 24B fits at Q8 in about 28.3 GB.

How many tokens per second does an RTX 5090 generate?

It depends entirely on the size of the model, because generating a token means reading every active weight out of memory. At 1,792 GB/s and 70% efficiency, this card produces an estimated 270 tokens per second on an 8B model at Q4, and about 66 on the largest model it holds, Qwen3 32B. Reading speed is roughly 5 to 8 words per second, so anything above 25 feels immediate.

Is 32 GB enough for running LLMs locally?

It runs 13 of the 16 models tracked here at Q4_K_M with 8k of context, up to 32.8B parameters. The honest test is not the model list but the context: Qwen3 32B on this card holds about 37,888 tokens before memory runs out.

RTX 5090 or RTX 4090 for local models?

This card holds more: 32 GB against 24 GB, so it runs 13 of these models to the RTX 4090's 13. On speed, this card is faster: 1,792 GB/s against 1,008 GB/s, and bandwidth is what sets chat speed.

Is the RTX 5090 worth it over a 4090 for AI?

For text generation, yes, more than the specification sheet suggests: 1,792 GB/s against 1,008 GB/s is roughly 78% more memory read per second, and that is what sets generation speed. The 8 GB of extra capacity matters less than the bandwidth, though it does bring 32B models within reach at longer context.

How much power does an RTX 5090 need for AI workloads?

The board is rated at 575 W, and inference holds it near that for as long as the model is generating. Plan the power supply and the room around it — in a small space this is a meaningful amount of heat, which is a real argument for renting the occasional large run rather than owning it.

  • Memory: weights (parameters × bytes per parameter) + KV cache (2 × layers × KV heads × head dim × context × bytes) + 0.5 GB runtime + 4% of weights. Full derivation in the VRAM calculator.
  • Speed: 70% × bandwidth ÷ active weight bytes, single stream, no batching, prompt processing excluded. See the tokens-per-second estimator.
  • Card specification: 32 GB, 1,792 GB/s, 575 W — manufacturer figures.
  • Model architecture values come from each model's published config; the self-hosted LLM guide explains what each one changes.