What LLMs can an H100 SXM run?

Which open models run on an H100 SXM: what fits at Q4 and Q8, estimated tokens per second, how much context fits, and where the card stops.

Updated 9 September 2026 · estimates are labelled as estimates

80 GB VRAM 3,352 GB/s 700 W 2022 Rented, not bought

H100 SXM 80 GB addresses 80 GB at 3,352 GB/s. Of the 16 open models tracked on this site, it runs 16 at Q4_K_M with 8k of context. The largest is gpt-oss 120B, needing about 71.7 GB and generating an estimated 793 tokens per second — faster than you can read.

The fastest memory here by a wide margin, and almost always more than one person needs.

What an H100 SXM runs, model by model

Every model at Q4_K_M with 8k of context, the setting most people actually use. This card has 76 GB to spend once the 5% safety margin comes off its 80 GB, and it reads that memory at 3,352 GB/s. Those two numbers decide the last two columns: the first says what loads, the second says how fast it answers. Both are estimates from stated formulas, and the VRAM calculator shows the working.

ModelSizeNeedsEst. tokens/sOn this card
Llama 3.2 3B 3.2B 3.4 GB 1264 fits
Llama 3.1 8B 8B 6.4 GB 506 fits
Qwen3 8B 8.2B 6.7 GB 493 fits
Gemma 3 12B 12.2B 11.1 GB 332 fits
Qwen3 14B 14.8B 10.8 GB 273 fits
Phi-4 14B 14.7B 11 GB 275 fits
gpt-oss 20B 21B · 3.6B active 13.6 GB 1124 fits
Mistral Small 3.1 24B 24B 16.3 GB 169 fits
Gemma 3 27B 27.4B 21.2 GB 148 fits
Qwen3 30B-A3B (MoE) 30.5B · 3.3B active 19.7 GB 1226 fits
Qwen3 32B 32.8B 22.4 GB 123 fits
Qwen2.5 Coder 32B 32.8B 22.4 GB 123 fits
DeepSeek-R1 Distill Qwen 32B 32.8B 22.4 GB 123 fits
Llama 3.3 70B 70.6B 45.8 GB 57 fits
DeepSeek-R1 Distill Llama 70B 70.6B 45.8 GB 57 fits
gpt-oss 120B 117B · 5.1B active 71.7 GB 793 tight fit

On this card that means anything at or under 76 GB counts as fitting, and anything above 68 GB is marked tight: it loads, but there is little left for the conversation, and a desktop on the same 80 GB can push it over.

The model to actually run on it

Llama 3.3 70B is the best use of this card: 45.8 GB of the 80 GB available, an estimated 57 tokens per second, comfortable for chat, and 100,352 tokens of context still available.

gpt-oss 120B is the largest model the card will hold, at 71.7 GB, but it is the wrong daily driver: filling 90% of memory with weights leaves room for only 66,560 tokens of conversation, and it generates at 793 tokens per second against 57. More parameters are not worth a context window that runs out mid-document.

If quality matters more than parameter count, Qwen3 32B fits at Q8_0 in about 38.8 GB. Q8 is near-lossless; a smaller model at Q8 often beats a larger one squeezed into Q3.

How much context actually fits

Model size is the question people ask; context is the one that bites. The KV cache grows linearly with the conversation, so a model that loads comfortably can still run out of memory halfway through a long document. On 80 GB, the context you get is whatever the weights leave behind — which is why two cards that run the same model can feel completely different in use. These are the largest contexts this card holds, at Q4_K_M weights.

ModelMax context, FP16 KVWith Q8 KV cache
gpt-oss 120B 66,560 tokens 131,072 tokens
Llama 3.3 70B 100,352 tokens 131,072 tokens
DeepSeek-R1 Distill Llama 70B 100,352 tokens 131,072 tokens
Qwen3 32B 131,072 tokens 131,072 tokens

Quantising the KV cache to Q8 roughly doubles what you can hold — gpt-oss 120B on this card goes from 66,560 to 131,072 tokens — at a quality cost most people never notice. Capped at 128k tokens here; a model's own architecture may stop lower, and sliding-window models such as Gemma 3 use less KV than the formula assumes.

Where this card stops

Nothing in this list is out of reach. Every model tracked here fits at Q4_K_M with 8k of context, so the ceiling you will meet is context length and batch size, not parameter count.

The honest take

At 3,352 GB/s the H100 SXM generates text faster than anything else on this page — roughly three and a half times a 4090 on the same model. That speed is designed to be amortised across many simultaneous users, which is why it is the backbone of inference providers and rarely the right answer for a single conversation. For one person chatting, most of what you rent goes unused.

What it is good at

  • 3,352 GB/s of HBM3: even a 70B model answers faster than you can read.
  • FP8 support, which newer runtimes use to cut memory and increase throughput.
  • Where batching matters, nothing else on this list is close.

What it is not

  • 700 W and an SXM baseboard — it does not go in a PCIe slot at all.
  • The most expensive class to rent, for speed a single user cannot consume.
  • 80 GB is less memory than an RTX PRO 6000, despite costing far more.

The thing people get wrong: The SXM version is not a card you can install. It mounts on an OEM baseboard, so "buying an H100" means buying a server. The PCIe variant exists and is meaningfully slower.

Buy or rent

Rent, and only when serving many users at once or running something genuinely time-critical. For one person, an A100 or a 48 GB card delivers the same experience for much less.

There is no purchase price to reason about, only an hourly one, and it differs by provider and by week. Put the figure you are actually looking at, with your electricity rate and your real monthly hours, into the build vs rent calculator — it already knows this card draws 700 W under load. It returns the month owning becomes cheaper, or tells you it never does. Hours per month decides it far more often than the price does.

Nodegrove is the rented side of that comparison: a workspace that stays saved, with a GPU attached only while you are using it. It is the right answer when your hours are low or your needs change; buying is the right answer when they are high and stable.

Compared with the alternatives

  • A10080 GB · 2,039 GB/s · best fit Llama 3.3 70B at ~35 tok/s. The old datacenter workhorse, still fast where it counts, and cheap to rent.
  • RTX PRO 6000 Blackwell96 GB · 1,792 GB/s · best fit gpt-oss 120B at ~424 tok/s. Ninety-six gigabytes and 5090-class bandwidth on one card.
  • L40S48 GB · 864 GB/s · best fit Qwen3 32B at ~32 tok/s. The 48 GB card you meet as a rental line item, not as a purchase.

All 13 cards are compared side by side on the GPU index, and the same numbers from the model's point of view are on each model fit table. To check one specific pairing rather than read a table, use the can-I-run-it checker.

Questions

Can an H100 SXM run a 70B model?

Yes. Llama 3.3 70B at Q4_K_M with 8k of context needs about 45.8 GB, inside the 80 GB this card can address, and it generates an estimated 57 tokens per second — comfortable for chat.

What is the best model to run on an H100 SXM?

Llama 3.3 70B. At Q4_K_M and 8k context it needs about 45.8 GB of the 80 GB available and generates an estimated 57 tokens per second, which is comfortable for chat. If quality matters more than size, Qwen3 32B fits at Q8 in about 38.8 GB.

How many tokens per second does an H100 SXM generate?

It depends entirely on the size of the model, because generating a token means reading every active weight out of memory. At 3,352 GB/s and 70% efficiency, this card produces an estimated 506 tokens per second on an 8B model at Q4, and about 793 on the largest model it holds, gpt-oss 120B. Reading speed is roughly 5 to 8 words per second, so anything above 25 feels immediate.

Is 80 GB enough for running LLMs locally?

It runs 16 of the 16 models tracked here at Q4_K_M with 8k of context, up to 117B parameters. The honest test is not the model list but the context: gpt-oss 120B on this card holds about 66,560 tokens before memory runs out.

H100 SXM or A100 for local models?

Both address about the same memory, so they run the same models. On speed, this card is faster: 3,352 GB/s against 2,039 GB/s, and bandwidth is what sets chat speed.

Can I buy an H100 for a workstation?

Not the SXM version, which is what most H100 references mean. It mounts on an OEM baseboard rather than a PCIe slot, so buying one means buying a server built around it, at 700 W per card. A PCIe variant exists and is meaningfully slower.

Is an H100 overkill for one person?

Almost always. Its 3,352 GB/s is designed to be shared across many simultaneous requests, and a single conversation cannot consume that. For one person chatting with a large model, an A100 or a 48 GB card delivers a very similar experience for a fraction of the rental cost.

  • Memory: weights (parameters × bytes per parameter) + KV cache (2 × layers × KV heads × head dim × context × bytes) + 0.5 GB runtime + 4% of weights. Full derivation in the VRAM calculator.
  • Speed: 70% × bandwidth ÷ active weight bytes, single stream, no batching, prompt processing excluded. See the tokens-per-second estimator.
  • Card specification: 80 GB, 3,352 GB/s, 700 W — manufacturer figures.
  • Model architecture values come from each model's published config; the self-hosted LLM guide explains what each one changes.