Image and video models are measured differently from language models. There is no context to budget for: what decides whether a model runs is the size of the files ComfyUI loads, stage by stage, and how much memory is left for the work. Every figure below comes from those files. Of the 10 models here, 7 load ComfyUI's own template set with room to spare on a 24 GB card, and 9 do with the smallest file that card can use.
Every model, compared
Sorted by what ComfyUI's official template needs at its largest stage. "Smallest card" is the smallest card listed on this site that loads the template set with room to spare (under 85% of its memory).
| Model | Type | Template peak | Download | Smallest card | On 24 GB | Licence |
|---|---|---|---|---|---|---|
| Anima | image | 4.4 GB | 5.6 GB | RTX 4070 | BF16 | non-commercial |
| FLUX.2 · klein 4B | image | 8.0 GB | 12.5 GB | RTX 4070 | BF16 | non-commercial |
| Qwen-Image-2.1 | image | 9.5 GB | 26.8 GB | RTX 4070 | BF16 | non-commercial |
| Z-Image · Turbo | image | 12.7 GB | 20.7 GB | RTX 5060 Ti | BF16 | Apache 2.0 |
| Krea 2 · Turbo | image | 13.4 GB | 18.6 GB | RTX 5060 Ti | FP8 scaled | Krea 2 Community License |
| Wan 2.2 · Text to video, A14B | video | 14.5 GB | 35.6 GB | RTX 4090 | FP8 scaled | Apache 2.0 |
| HunyuanVideo 1.5 | video | 19.2 GB | 29.0 GB | RTX 4090 | FP16 | territory limits |
| LTX-2.5 · Distilled | video + audio | 24.3 GB | 44.9 GB | RTX 5090 | streams | LTX-2 Community License |
| MiniMax H3 | video + audio | 24.4 GB | 40.1 GB | RTX 5090 | W6A8, pruned | territory limits |
| Qwen-Image-Edit-2511 | image edit | 41.1 GB | 50.5 GB | H100 SXM | FP8 mixed | Apache 2.0 |
Check any model on your own card, or pick a precision by hand, in the ComfyUI VRAM checker. Wan 2.5, 2.6, 2.7 and 3.0 are missing on purpose: none has open weights (checked 2026-10-02).
The models
Anima
An anime image model from CircleStone Labs, made with Comfy Org and built on NVIDIA's Cosmos-Predict2 2B. Template peak 4.4 GB; the smallest file loads on a RTX 4070.
FLUX.2
Black Forest Labs' second FLUX generation: dev (November 2025), a 32-billion-parameter model, and klein (January 2026), 4B and 9B models built for consumer cards. Template peak 8.0 GB; the smallest file loads on a RTX 4070.
Qwen-Image-2.1
Qwen's image model from September 2026: a 7-billion-parameter generator with native 2K output, driven by an 8-billion-parameter Qwen3-VL text encoder. Template peak 9.5 GB; the smallest file loads on a RTX 4070.
Z-Image
Tongyi-MAI's 6-billion-parameter image model, one of the few here under Apache 2.0. Template peak 12.7 GB; the smallest file loads on a RTX 4070.
Krea 2
Krea's open image model from June 2026: a 12-billion-parameter transformer in two versions. Template peak 13.4 GB; the smallest file loads on a RTX 5060 Ti.
Wan 2.2
Alibaba's Wan 2.2, from July 2025, remains the newest Wan with open weights. Template peak 14.5 GB; the smallest file loads on a RTX 5060 Ti.
HunyuanVideo 1.5
Tencent's 8.3-billion-parameter video model from November 2025: 480p and 720p, 121 frames by default, with a separate super-resolution model that takes the result to 1080p. Template peak 19.2 GB; the smallest file loads on a RTX 4070.
LTX-2.5
Lightricks' open video model from August 2026, a 22-billion-parameter transformer that generates sound with the picture. Template peak 24.3 GB; the smallest file loads on a RTX 5090.
MiniMax H3
MiniMax's open video model from August 2026, and the only one here that makes the soundtrack with the picture: stereo audio, 4 to 15 seconds at 24 frames a second. Template peak 24.4 GB; the smallest file loads on a RTX 4090.
Qwen-Image-Edit-2511
Qwen's image-editing model from December 2025: give it an image and an instruction. Template peak 41.1 GB; the smallest file loads on a RTX 4090.
How the numbers work
A ComfyUI workflow runs in stages: the text encoder turns your prompt into embeddings, an optional prompt enhancer rewrites it, then the diffusion model samples and the VAE decodes. Models load when a node needs them and move out when another needs the room, so what has to fit is the largest stage, not the sum of every file. Wan 2.2's A14B models are two experts that take turns, so one is on the card at a time and both are downloaded.
What is left after the weights is working memory, and that is the part nobody publishes per model: it grows with resolution, frame count and batch size, and it depends on the attention kernel. So each model page quotes what the vendor says it needs, word for word, beside the file arithmetic, and says which figure measures what.
- File sizes: exact byte counts from the Hugging Face API, for Comfy-Org's repackaged repos and the vendors' own ComfyUI-ready files, read on 2 October 2026.
- Template sets: the official workflow guides on docs.comfy.org, linked from each model page.
- Card memory: manufacturers' specifications, as on each GPU page. NVIDIA cards only; ComfyUI also runs on AMD and Apple silicon, which we have not sized.