Qwen-Image-Edit-2511 is Qwen's open image-editing model. ComfyUI's own workflow downloads 50.5 GB and needs 41.1 GB on the card at its largest stage, the diffusion stage. The smallest card listed here that loads that set with room to spare is the H100 SXM. With a smaller file, the FP8 mixed version, it loads on an RTX 4090.
What Qwen-Image-Edit-2511 is
Qwen's image-editing model from December 2025: give it an image and an instruction. A 20-billion-parameter model under Apache 2.0, with the same Qwen2.5-VL text encoder as the original Qwen-Image.
20.4B parameters Text encoder: Qwen2.5-VL 7B Apache 2.0 December 2025
Which Qwen-Image-Edit-2511 file for your card
For each memory size, the highest-precision diffusion file that loads with room to spare (a file that only just fits is used when nothing else does), and what is left for working memory while it samples. "Streams" means even the smallest file is larger than the card: ComfyUI still runs it by streaming weights from system RAM, more slowly. FP4 files are offered only to the Blackwell cards that run them natively.
| Card | Usable | File that loads | Largest stage | Left for the work | Verdict |
|---|---|---|---|---|---|
| RTX 4070 | 12 GB | smallest: FP8 mixed | 20.8 GB | – | Streams |
| RTX 5060 Ti | 16 GB | smallest: FP8 mixed | 20.8 GB | – | Streams |
| RTX 4090 | 24 GB | FP8 mixed | 20.8 GB | 3.2 GB | Loads, tight |
| RTX 5090 | 32 GB | FP8 mixed | 20.8 GB | 11.2 GB | Loads fully |
| RTX 6000 Ada | 48 GB | FP8 mixed | 20.8 GB | 27.2 GB | Loads fully |
| H100 SXM | 80 GB | BF16 | 41.1 GB | 38.9 GB | Loads fully |
| RTX PRO 6000 Blackwell | 96 GB | BF16 | 41.1 GB | 54.9 GB | Loads fully |
| DGX Spark | 126 GB | BF16 | 41.1 GB | 84.9 GB | Loads fully |
ComfyUI's template: BF16 diffusion model, FP8 scaled text encoder. Stages: text encoder 9.4 GB, diffusion 41.1 GB. Download 50.5 GB.
Check another card, or pick a file by hand, in the ComfyUI VRAM checker.
Every Qwen-Image-Edit-2511 file ComfyUI loads
Sizes are the exact byte counts Hugging Face reports. Files marked "template" are the ones ComfyUI's official workflow downloads.
| Part | Precision | Size | File |
|---|---|---|---|
| Diffusion model | BF16 · template | 40.86 GB | qwen_image_edit_2511_bf16.safetensors |
| Diffusion model | FP8 mixed | 20.53 GB | qwen_image_edit_2511_fp8mixed.safetensors |
| Diffusion model | INT8 ConvRot | 20.50 GB | qwen_image_edit_2511_int8_convrot.safetensors |
| Text encoder | BF16 | 16.58 GB | qwen_2.5_vl_7b.safetensors |
| Text encoder | FP8 scaled · template | 9.38 GB | qwen_2.5_vl_7b_fp8_scaled.safetensors |
| Text encoder | NVFP4 | 6.11 GB | qwen_2.5_vl_7b_nvfp4.safetensors |
| VAE | BF16 · template | 0.25 GB | qwen_image_vae.safetensors |
What Qwen says it needs
- "low-GPU-memory layer-by-layer offload (inference within 4GB VRAM)" (DiffSynth-Studio's layer-by-layer offloading, not ComfyUI, source)
Those figures and this page measure different things. A vendor's figure covers its own script at its default resolution, working memory included. The tables here count the weights each stage holds, which is exact, and leave the working memory as the space that remains, because it depends on your resolution and attention kernel and nobody publishes it per model. If generation runs out of memory with a file that loads, lower the resolution before stepping down a precision. ComfyUI's README says it "can run even the biggest open source models on as low as 4GB vram + 8GB ram" by streaming weights (source), so "streams" means slower, not impossible.
LoRA training for Qwen-Image-Edit-2511
Only configurations a trainer's own documentation gives a VRAM figure for. Settings change the figure a great deal, so read the setting column before trusting the number.
| Trainer | VRAM | Setting | Source |
|---|---|---|---|
| musubi-tuner | 42 GB | 1024×1024, batch 1, no quantisation; measured on Qwen-Image, and Edit needs more for its control images | docs |
| musubi-tuner | 24 GB | as above with an FP8 model and 16 blocks swapped | docs |
| musubi-tuner | 12 GB | as above with 45 blocks swapped (64 GB of system RAM recommended) | docs |
What Qwen-Image-Edit-2511 is good at, and what to watch
Good at
- Instruction-based edits: material swaps, object changes, restyling.
- Commercial use under Apache 2.0.
- LoRA training on cards from 12 GB up, with musubi-tuner's documented settings.
Watch out for
- ComfyUI's template loads the BF16 model, although FP8 and INT8 versions at half the size exist.
- Qwen's "within 4GB VRAM" figure is for DiffSynth-Studio's layer-by-layer offloading, not ComfyUI.
Or skip the card
ComfyUI's Qwen-Image-Edit-2511 workflow peaks at 41.1 GB, more than a 24 GB card holds with room to work. A Nodegrove workspace attaches a 96 GB GPU and keeps your ComfyUI setup, models and outputs saved between sessions. In early access, not open yet.
Questions
How much VRAM does Qwen-Image-Edit-2511 need?
In ComfyUI's own workflow, the largest stage is the diffusion stage at 41.1 GB, and the whole workflow downloads 50.5 GB. Working memory for generation comes on top and grows with resolution. Qwen itself says: "low-GPU-memory layer-by-layer offload (inference within 4GB VRAM)" (DiffSynth-Studio's layer-by-layer offloading, not ComfyUI).
Can I run Qwen-Image-Edit-2511 on 8 GB of VRAM?
Not fully. The smallest file here, FP8 mixed at 20.5 GB, needs 20.8 GB at its largest stage, more than 8 GB holds. ComfyUI will still run it by streaming weights from system RAM, which works but is slower.
Can I run Qwen-Image-Edit-2511 on 12 GB of VRAM?
Not fully. The smallest file here, FP8 mixed at 20.5 GB, needs 20.8 GB at its largest stage, more than 12 GB holds. ComfyUI will still run it by streaming weights from system RAM, which works but is slower.
Can I run Qwen-Image-Edit-2511 on 16 GB of VRAM?
Not fully. The smallest file here, FP8 mixed at 20.5 GB, needs 20.8 GB at its largest stage, more than 16 GB holds. ComfyUI will still run it by streaming weights from system RAM, which works but is slower.
Which Qwen-Image-Edit-2511 file should I download?
The highest precision that loads fully on your card. On a 24 GB card that is the FP8 mixed file (20.5 GB). ComfyUI's template downloads BF16 with the FP8 scaled text encoder. FP4 files (NVFP4) are made for Blackwell cards (RTX 50, RTX PRO 6000, DGX Spark); on older cards, use FP8 or INT8.
- Files and sizes: Comfy-Org/Qwen-Image-Edit_ComfyUI on Hugging Face, read from the Hugging Face API. Template files: ComfyUI's Qwen-Image-Edit-2511 guide.
- Model and licence: Qwen/Qwen-Image-Edit-2511 and the Apache 2.0.
- Card memory: manufacturers' specifications, as on each GPU page. The checker covers NVIDIA cards; ComfyUI also runs on AMD and Apple silicon, where file formats and speed differ and we have not sized them.
- Fit rule: a stage loads fully under 85% of usable memory and is tight up to 95%, the same margins the LLM pages use.