Qwen-Image-2.1 is Qwen's open image model. ComfyUI's own workflow downloads 26.8 GB and needs 9.5 GB on the card at its largest stage, the prompt enhancer. The smallest card listed here that loads that set with room to spare is the RTX 4070.
What Qwen-Image-2.1 is
Qwen's image model from September 2026: a 7-billion-parameter generator with native 2K output, driven by an 8-billion-parameter Qwen3-VL text encoder. Its text side, the encoder plus the 9-billion-parameter prompt rewriter in ComfyUI's template, is larger than the image generator, and that is what decides the memory.
7.1B parameters Text encoder: Qwen3-VL 8B Qwen Research License September 2026
Before you download
Research and evaluation only; commercial use needs a separate licence from Qwen. Read the Qwen Research License itself before using it for work.
Which Qwen-Image-2.1 file for your card
For each memory size, the highest-precision diffusion file that loads with room to spare (a file that only just fits is used when nothing else does), and what is left for working memory while it samples. "Streams" means even the smallest file is larger than the card: ComfyUI still runs it by streaming weights from system RAM, more slowly. FP4 files are offered only to the Blackwell cards that run them natively.
| Card | Usable | File that loads | Largest stage | Left for the work | Verdict |
|---|---|---|---|---|---|
| RTX 4070 | 12 GB | INT8 ConvRot | 9.5 GB | 4.1 GB | Loads fully |
| RTX 5060 Ti | 16 GB | INT8 ConvRot | 9.5 GB | 8.1 GB | Loads fully |
| RTX 4090 | 24 GB | BF16 | 17.5 GB | 9.1 GB | Loads fully |
| RTX 5090 | 32 GB | BF16 | 17.5 GB | 17.1 GB | Loads fully |
| RTX 6000 Ada | 48 GB | BF16 | 17.5 GB | 33.1 GB | Loads fully |
| H100 SXM | 80 GB | BF16 | 17.5 GB | 65.1 GB | Loads fully |
| RTX PRO 6000 Blackwell | 96 GB | BF16 | 17.5 GB | 81.1 GB | Loads fully |
| DGX Spark | 126 GB | BF16 | 17.5 GB | 111.1 GB | Loads fully |
ComfyUI's template: INT8 ConvRot diffusion model, INT8 ConvRot text encoder. Stages: text encoder 9.4 GB, prompt enhancer 9.5 GB, diffusion 7.9 GB. Download 26.8 GB.
Check another card, or pick a file by hand, in the ComfyUI VRAM checker.
Every Qwen-Image-2.1 file ComfyUI loads
Sizes are the exact byte counts Hugging Face reports. Files marked "template" are the ones ComfyUI's official workflow downloads.
| Part | Precision | Size | File |
|---|---|---|---|
| Diffusion model | BF16 | 14.23 GB | qwen_image_2.1_bf16.safetensors |
| Diffusion model | INT8 ConvRot · template | 7.26 GB | qwen_image_2.1_int8_convrot.safetensors |
| Text encoder | BF16 | 17.53 GB | qwen3vl_8b_bf16.safetensors |
| Text encoder | INT8 ConvRot · template | 9.35 GB | qwen3vl_8b_int8_convrot.safetensors |
| Text encoder | W4A8 | 6.31 GB | qwen3vl_8b_w4a8.safetensors |
| Prompt enhancer | INT8 ConvRot · template | 9.47 GB | qwen3.5_9b_qwen_image_2.1_pe_t2i.int8_convrot.safetensors |
| VAE | BF16 · template | 0.68 GB | qwen_image_2.1_vae_bf16.safetensors |
What Qwen says it needs
Qwen states no VRAM figure on the model card or in its GitHub README. The closest it comes: "With just 7B parameters in its visual generation component" (source).
Those figures and this page measure different things. A vendor's figure covers its own script at its default resolution, working memory included. The tables here count the weights each stage holds, which is exact, and leave the working memory as the space that remains, because it depends on your resolution and attention kernel and nobody publishes it per model. If generation runs out of memory with a file that loads, lower the resolution before stepping down a precision. ComfyUI's README says it "can run even the biggest open source models on as low as 4GB vram + 8GB ram" by streaming weights (source), so "streams" means slower, not impossible.
LoRA training for Qwen-Image-2.1
No trainer we checked publishes a VRAM figure for training Qwen-Image-2.1 yet (kohya's sd-scripts and musubi-tuner, ostris' ai-toolkit, diffusion-pipe, and Qwen's own repo). When one does, it goes here.
What Qwen-Image-2.1 is good at, and what to watch
Good at
- Native 2K images.
- The same files run ComfyUI's image-edit template as well as text to image.
- INT8 files at about half the size of BF16.
Watch out for
- The Qwen Research License allows research and evaluation only. Commercial use needs a separate licence from Qwen.
- The prompt rewriter is a separate language model that ComfyUI's template loads by default.
- Qwen publishes no VRAM figure.
Questions
How much VRAM does Qwen-Image-2.1 need?
In ComfyUI's own workflow, the largest stage is the prompt enhancer at 9.5 GB, and the whole workflow downloads 26.8 GB. Working memory for generation comes on top and grows with resolution. Qwen publishes no VRAM figure.
Can I run Qwen-Image-2.1 on 8 GB of VRAM?
Not fully. The smallest file here, INT8 ConvRot at 7.3 GB, needs 9.5 GB at its largest stage, more than 8 GB holds. ComfyUI will still run it by streaming weights from system RAM, which works but is slower.
Can I run Qwen-Image-2.1 on 12 GB of VRAM?
Yes, with the INT8 ConvRot file: its largest stage, the prompt enhancer, is 9.5 GB, which loads on 12 GB with 4.1 GB left for working memory.
Can I run Qwen-Image-2.1 on 16 GB of VRAM?
Yes, with the INT8 ConvRot file: its largest stage, the prompt enhancer, is 9.5 GB, which loads on 16 GB with 8.1 GB left for working memory.
Which Qwen-Image-2.1 file should I download?
The highest precision that loads fully on your card. On a 24 GB card that is the BF16 file (14.2 GB). ComfyUI's template downloads INT8 ConvRot with the INT8 ConvRot text encoder. FP4 files (NVFP4) are made for Blackwell cards (RTX 50, RTX PRO 6000, DGX Spark); on older cards, use FP8 or INT8.
- Files and sizes: Comfy-Org/Qwen-Image-2.1 on Hugging Face, read from the Hugging Face API. Template files: ComfyUI's Qwen-Image-2.1 guide.
- Model and licence: Qwen/Qwen-Image-2.1 and the Qwen Research License.
- Card memory: manufacturers' specifications, as on each GPU page. The checker covers NVIDIA cards; ComfyUI also runs on AMD and Apple silicon, where file formats and speed differ and we have not sized them.
- Fit rule: a stage loads fully under 85% of usable memory and is tight up to 95%, the same margins the LLM pages use.