ComfyUI VRAM checker: which model file fits your GPU

Pick an open image or video model and your GPU: which precision to download, what each stage needs, and what is left for generating. Exact file sizes.

Updated · estimates are labelled as estimates
On this page
ComfyUI VRAM checker
An image or video model on your card: which file loads, and what is left for the work
Model
Your GPU
– –
Largest stage–
Card has–
Left for the work–
Download–
Every diffusion file on this card
PrecisionFileLargest stageVerdict

How the answer is worked out

A ComfyUI workflow runs in stages, and each stage loads only what it needs:

  • The text encoder turns your prompt into embeddings. For several current models it is a full language model of its own, sometimes larger than the image model.
  • A prompt enhancer, in some official templates, rewrites the prompt first. It is a separate model with its own stage.
  • The diffusion model samples the image or video, with the VAE alongside to decode it. Wan 2.2's A14B models are two experts that take turns, so one is on the card at a time.

The largest of those stages is what has to fit. The checker finds, for your card, the highest-precision diffusion file whose largest stage loads fully, picks the best text encoder that also fits, and shows what is left for working memory. File sizes are the exact byte counts Hugging Face reports for the 97 files ComfyUI loads across these models. Files in Blackwell formats (NVFP4, MXFP8) are offered only to Blackwell cards: RTX 50, RTX PRO 6000 and DGX Spark.

What does not fit is not refused: ComfyUI streams weights from system RAM, according to its README. "Streams" in the checker means slower, not impossible.

Every model, one page each

Each model has a page with its full file list, a card-by-card table, what its maker says it needs, documented LoRA training settings, and its licence terms.

Or see them side by side: every image and video model compared. For language models, the LLM checker does the same job.

Questions

How do I know which ComfyUI model file fits my GPU?

Look at the largest stage, not the sum of the files. ComfyUI loads the text encoder, runs it, and moves it out when the diffusion model needs the room, so what has to fit is the bigger of the two, plus the VAE with the diffusion model. A file loads fully when that stage stays under 85% of your card's memory, and tightly up to 95%. The checker does that for every file of every model it lists.

What happens if the model does not fit in VRAM?

ComfyUI streams the part that does not fit from system RAM instead of refusing to run. Its README says it can run even the biggest open models on as low as 4 GB of VRAM with 8 GB of RAM this way. It works and it is slower; how much slower depends on your PCIe link and the resolution, so we do not put a number on it.

Can I run Wan 2.2 on a 12 GB card?

Not fully. Even the smallest text-to-video expert, FP8 scaled, needs 14.5 GB at its largest stage, more than a 12 GB card holds, so ComfyUI streams part of it from system RAM. On a 24 GB card the FP8 scaled file loads fully.

Why not just add up the file sizes?

Because the files are not on the card at the same time, and because two of them may never be: Wan 2.2's text-to-video template downloads 35.6 GB, two experts and a text encoder, but its largest stage is 14.5 GB. Adding everything up overstates what you need; looking only at the diffusion file understates it when the text encoder is the larger one, as it is for some models here.

Does the checker estimate working memory?

No. Working memory grows with resolution, frame count and batch size, depends on the attention kernel, and no vendor publishes it per model. The checker shows what is left after the weights and quotes what each vendor says it needs, so you can see both and judge.

  • Files and sizes: Comfy-Org's repackaged repositories and the vendors' own ComfyUI-ready files on Hugging Face, exact byte counts read from the Hugging Face API on 2 October 2026. Template sets: the official guides on docs.comfy.org.
  • Card memory: manufacturers' specifications. NVIDIA cards only; ComfyUI also runs on AMD and Apple silicon, where formats and speed differ.
  • Fit rule: a stage loads fully under 85% of usable memory and is tight up to 95%, the same margins the LLM tools use.