What LLMs can an RTX 6000 Ada run?

Which open models run on an RTX 6000 Ada: what fits at Q4 and Q8, estimated tokens per second, how much context fits, and where the card stops.

Updated 9 September 2026 · estimates are labelled as estimates

48 GB VRAM 960 GB/s 300 W 2022 Workstation card

RTX 6000 Ada 48 GB addresses 48 GB at 960 GB/s. Of the 16 open models tracked on this site, it runs 13 at Q4_K_M with 8k of context. The largest is Qwen3 32B, needing about 22.4 GB and generating an estimated 35 tokens per second — comfortable for chat.

Forty-eight gigabytes in a normal computer, at 300 watts, without the noise.

What an RTX 6000 Ada runs, model by model

Every model at Q4_K_M with 8k of context, the setting most people actually use. This card has 45.6 GB to spend once the 5% safety margin comes off its 48 GB, and it reads that memory at 960 GB/s. Those two numbers decide the last two columns: the first says what loads, the second says how fast it answers. Both are estimates from stated formulas, and the VRAM calculator shows the working.

ModelSizeNeedsEst. tokens/sOn this card
Llama 3.2 3B 3.2B 3.4 GB 362 fits
Llama 3.1 8B 8B 6.4 GB 145 fits
Qwen3 8B 8.2B 6.7 GB 141 fits
Gemma 3 12B 12.2B 11.1 GB 95 fits
Qwen3 14B 14.8B 10.8 GB 78 fits
Phi-4 14B 14.7B 11 GB 79 fits
gpt-oss 20B 21B · 3.6B active 13.6 GB 322 fits
Mistral Small 3.1 24B 24B 16.3 GB 48 fits
Gemma 3 27B 27.4B 21.2 GB 42 fits
Qwen3 30B-A3B (MoE) 30.5B · 3.3B active 19.7 GB 351 fits
Qwen3 32B 32.8B 22.4 GB 35 fits
Qwen2.5 Coder 32B 32.8B 22.4 GB 35 fits
DeepSeek-R1 Distill Qwen 32B 32.8B 22.4 GB 35 fits
Llama 3.3 70B 70.6B 45.8 GB no
DeepSeek-R1 Distill Llama 70B 70.6B 45.8 GB no
gpt-oss 120B 117B · 5.1B active 71.7 GB no

On this card that means anything at or under 45.6 GB counts as fitting, and anything above 40.8 GB is marked tight: it loads, but there is little left for the conversation, and a desktop on the same 48 GB can push it over.

The model to actually run on it

Qwen3 32B is the best use of this card: 22.4 GB of the 48 GB available, an estimated 35 tokens per second, comfortable for chat, and 96,256 tokens of context still available.

It also fits at Q8_0, in about 38.8 GB, which is worth taking whenever the memory allows: Q8 is near-lossless where Q4 costs a little accuracy.

How much context actually fits

Model size is the question people ask; context is the one that bites. The KV cache grows linearly with the conversation, so a model that loads comfortably can still run out of memory halfway through a long document. On 48 GB, the context you get is whatever the weights leave behind — which is why two cards that run the same model can feel completely different in use. These are the largest contexts this card holds, at Q4_K_M weights.

ModelMax context, FP16 KVWith Q8 KV cache
Qwen3 32B 96,256 tokens 131,072 tokens
Qwen2.5 Coder 32B 96,256 tokens 131,072 tokens
DeepSeek-R1 Distill Qwen 32B 96,256 tokens 131,072 tokens
Qwen3 30B-A3B (MoE) 131,072 tokens 131,072 tokens

Quantising the KV cache to Q8 roughly doubles what you can hold — Qwen3 32B on this card goes from 96,256 to 131,072 tokens — at a quality cost most people never notice. Capped at 128k tokens here; a model's own architecture may stop lower, and sliding-window models such as Gemma 3 use less KV than the formula assumes.

Where this card stops

The first model out of reach is Llama 3.3 70B: about 45.8 GB at Q4_K_M and 8k context, against 48 GB of usable memory. That is close enough to be worth a caveat: it is inside the raw 48 GB and only misses the 95% line this site uses for a fit. On a card with a display attached, treat it as no. On a headless card, with a shorter context, it runs. Dropping to Q3_K_M would need about 37.7 GB, which fits, though Q3 loses enough quality that a smaller model at Q4 is usually the better trade. Offloading the remainder to system RAM works and is ten to fifty times slower; it is a way to see a model run, not a way to use one.

The honest take

This is what you buy when 24 GB is genuinely not enough and the machine has to live under a desk. Forty-eight gigabytes with ECC, in a two-slot blower at 300 W, is a combination no GeForce card offers. Bandwidth is 960 GB/s — between a 3090 and a 4090 — so it is not fast for its class, it is capacious and civilised.

What it is good at

  • 48 GB is the tier where 70B models come into range. Llama 3.3 70B at Q4 lands within a percent of the memory budget, so it runs on a card with no display attached, with very little context to spare.
  • 300 W and a two-slot blower: it fits in a workstation and exhausts its own heat.
  • ECC memory and certified drivers, which matter for long fine-tuning runs.

What it is not

  • 960 GB/s is lower than a 4090 and far below a 5090. On models that fit both, a 4090 is faster.
  • Workstation pricing is several consumer cards for one card's worth of speed.
  • Two 3090s reach the same 48 GB for far less, if you can live with the heat, the noise and the split.

The thing people get wrong: It is slower than a 4090 at generating text. You are not buying performance, you are buying a 48 GB address space in a quiet 300 W package — which is exactly the right trade for some people and a waste of money for everyone else.

Buy or rent

Buy it when the constraint is physical: one machine, under a desk, quiet, 48 GB, no server rack. If the constraint is only capability, renting a 48 GB card for the hours you need it is far cheaper.

Workstation cards are sold on quotes rather than shelf prices, so there is no single number to publish. Put the figure you are actually looking at, with your electricity rate and your real monthly hours, into the build vs rent calculator — it already knows this card draws 300 W under load. It returns the month owning becomes cheaper, or tells you it never does. Hours per month decides it far more often than the price does.

Nodegrove is the rented side of that comparison: a workspace that stays saved, with a GPU attached only while you are using it. It is the right answer when your hours are low or your needs change; buying is the right answer when they are high and stable.

Compared with the alternatives

  • L40S48 GB · 864 GB/s · best fit Qwen3 32B at ~32 tok/s. The 48 GB card you meet as a rental line item, not as a purchase.
  • RTX PRO 6000 Blackwell96 GB · 1,792 GB/s · best fit gpt-oss 120B at ~424 tok/s. Ninety-six gigabytes and 5090-class bandwidth on one card.
  • RTX 409024 GB · 1,008 GB/s · best fit Qwen3 30B-A3B (MoE) at ~369 tok/s. The card most local-model advice is implicitly written for.

All 13 cards are compared side by side on the GPU index, and the same numbers from the model's point of view are on each model fit table. To check one specific pairing rather than read a table, use the can-I-run-it checker.

Questions

Can an RTX 6000 Ada run a 70B model?

Not at Q4_K_M in one card. Llama 3.3 70B needs about 45.8 GB at 8k context and this card can address 48 GB. Your options are a smaller model, a harsher quantisation with very little context, splitting across two cards, or renting a larger one for the hours you need it.

What is the best model to run on an RTX 6000 Ada?

Qwen3 32B. At Q4_K_M and 8k context it needs about 22.4 GB of the 48 GB available and generates an estimated 35 tokens per second, which is comfortable for chat. It also fits at Q8, at about 38.8 GB, which is worth taking when it fits.

How many tokens per second does an RTX 6000 Ada generate?

It depends entirely on the size of the model, because generating a token means reading every active weight out of memory. At 960 GB/s and 70% efficiency, this card produces an estimated 145 tokens per second on an 8B model at Q4, and about 35 on the largest model it holds, Qwen3 32B. Reading speed is roughly 5 to 8 words per second, so anything above 25 feels immediate.

Is 48 GB enough for running LLMs locally?

It runs 13 of the 16 models tracked here at Q4_K_M with 8k of context, up to 32.8B parameters. The honest test is not the model list but the context: Qwen3 32B on this card holds about 96,256 tokens before memory runs out.

RTX 6000 Ada or L40S for local models?

Both address about the same memory, so they run the same models. On speed, this card is faster: 960 GB/s against 864 GB/s, and bandwidth is what sets chat speed.

RTX 6000 Ada or two RTX 3090s for 48 GB?

Two 3090s reach the same 48 GB for far less money, and can pool it over NVLink. The RTX 6000 Ada gives you that memory in one card, at 300 W, in a two-slot blower that exhausts its own heat, with ECC and certified drivers. You are buying quiet, simple and supported rather than fast — a single 4090 generates text more quickly than this card does.

Is the RTX 6000 Ada good for fine-tuning?

It is one of the better single-card options for it: 48 GB with ECC holds training states that consumer cards cannot, and it runs at a power and noise level a desk can live with for the length of a long run. For a run you do once, renting the same capacity is far cheaper.

  • Memory: weights (parameters × bytes per parameter) + KV cache (2 × layers × KV heads × head dim × context × bytes) + 0.5 GB runtime + 4% of weights. Full derivation in the VRAM calculator.
  • Speed: 70% × bandwidth ÷ active weight bytes, single stream, no batching, prompt processing excluded. See the tokens-per-second estimator.
  • Card specification: 48 GB, 960 GB/s, 300 W — manufacturer figures.
  • Model architecture values come from each model's published config; the self-hosted LLM guide explains what each one changes.