Z-Image is Tongyi-MAI's open image model. ComfyUI's own workflow for Turbo downloads 20.7 GB and needs 12.7 GB on the card at its largest stage, the diffusion stage. The smallest card listed here that loads that set with room to spare is the RTX 5060 Ti. With a smaller file, the INT8 ConvRot version, it loads on an RTX 4070.
What Z-Image is
Tongyi-MAI's 6-billion-parameter image model, one of the few here under Apache 2.0. Turbo (November 2025) generates in 8 steps; Base (January 2026) is the undistilled model for fine-tuning and longer sampling.
Turbo: 6.2B · Base: 6.2B Text encoder: Qwen3 4B Apache 2.0 November 2025
Which Z-Image file for your card
For each memory size, the highest-precision diffusion file that loads with room to spare (a file that only just fits is used when nothing else does), and what is left for working memory while it samples. "Streams" means even the smallest file is larger than the card: ComfyUI still runs it by streaming weights from system RAM, more slowly. FP4 files are offered only to the Blackwell cards that run them natively.
Turbo
| Card | Usable | File that loads | Largest stage | Left for the work | Verdict |
|---|---|---|---|---|---|
| RTX 4070 | 12 GB | INT8 ConvRot | 8.0 GB | 5.5 GB | Loads fully |
| RTX 5060 Ti | 16 GB | BF16 | 12.7 GB | 3.3 GB | Loads fully |
| RTX 4090 | 24 GB | BF16 | 12.7 GB | 11.3 GB | Loads fully |
| RTX 5090 | 32 GB | BF16 | 12.7 GB | 19.3 GB | Loads fully |
| RTX 6000 Ada | 48 GB | BF16 | 12.7 GB | 35.3 GB | Loads fully |
| H100 SXM | 80 GB | BF16 | 12.7 GB | 67.3 GB | Loads fully |
| RTX PRO 6000 Blackwell | 96 GB | BF16 | 12.7 GB | 83.3 GB | Loads fully |
| DGX Spark | 126 GB | BF16 | 12.7 GB | 113.3 GB | Loads fully |
ComfyUI's template for Turbo: BF16 diffusion model, BF16 text encoder. Stages: text encoder 8.0 GB, diffusion 12.7 GB. Download 20.7 GB.
Base
| Card | Usable | File that loads | Largest stage | Left for the work | Verdict |
|---|---|---|---|---|---|
| RTX 4070 | 12 GB | INT8 ConvRot | 8.0 GB | 5.5 GB | Loads fully |
| RTX 5060 Ti | 16 GB | BF16 | 12.7 GB | 3.3 GB | Loads fully |
| RTX 4090 | 24 GB | BF16 | 12.7 GB | 11.3 GB | Loads fully |
| RTX 5090 | 32 GB | BF16 | 12.7 GB | 19.3 GB | Loads fully |
| RTX 6000 Ada | 48 GB | BF16 | 12.7 GB | 35.3 GB | Loads fully |
| H100 SXM | 80 GB | BF16 | 12.7 GB | 67.3 GB | Loads fully |
| RTX PRO 6000 Blackwell | 96 GB | BF16 | 12.7 GB | 83.3 GB | Loads fully |
| DGX Spark | 126 GB | BF16 | 12.7 GB | 113.3 GB | Loads fully |
ComfyUI's template for Base: BF16 diffusion model, BF16 text encoder. Stages: text encoder 8.0 GB, diffusion 12.7 GB. Download 20.7 GB.
Check another card, or pick a file by hand, in the ComfyUI VRAM checker.
Every Z-Image file ComfyUI loads
Sizes are the exact byte counts Hugging Face reports. Files marked "template" are the ones ComfyUI's official workflow downloads.
| Part | Precision | Size | Version | File |
|---|---|---|---|---|
| Diffusion model | BF16 · template | 12.31 GB | Turbo | z_image_turbo_bf16.safetensors |
| Diffusion model | INT8 ConvRot | 6.20 GB | Turbo | z_image_turbo_int8_convrot.safetensors |
| Diffusion model | NVFP4 | 4.51 GB | Turbo | z_image_turbo_nvfp4.safetensors |
| Diffusion model | BF16 | 12.31 GB | Base | z_image_bf16.safetensors |
| Diffusion model | INT8 ConvRot | 6.20 GB | Base | z_image_int8_convrot.safetensors |
| Text encoder | BF16 · template | 8.04 GB | all | qwen_3_4b.safetensors |
| Text encoder | FP8 mixed | 5.63 GB | all | qwen_3_4b_fp8_mixed.safetensors |
| Text encoder | FP4 mixed | 3.48 GB | all | qwen_3_4b_fp4_mixed.safetensors |
| VAE | FP32 · template | 0.34 GB | all | ae.safetensors |
What Tongyi-MAI says it needs
- "fits comfortably within 16G VRAM consumer devices" (Z-Image-Turbo, the vendor's own inference code, source)
Those figures and this page measure different things. A vendor's figure covers its own script at its default resolution, working memory included. The tables here count the weights each stage holds, which is exact, and leave the working memory as the space that remains, because it depends on your resolution and attention kernel and nobody publishes it per model. If generation runs out of memory with a file that loads, lower the resolution before stepping down a precision. ComfyUI's README says it "can run even the biggest open source models on as low as 4GB vram + 8GB ram" by streaming weights (source), so "streams" means slower, not impossible.
LoRA training for Z-Image
No trainer we checked publishes a VRAM figure for training Z-Image yet (kohya's sd-scripts and musubi-tuner, ostris' ai-toolkit, diffusion-pipe, and Tongyi-MAI's own repo). When one does, it goes here.
What Z-Image is good at, and what to watch
Good at
- Fast drafts: Turbo needs 8 steps where Base takes 28 to 50.
- Commercial use without conditions beyond the Apache notice.
- Images up to 2048 by 2048 pixels (Base).
Watch out for
- Tongyi-MAI's "fits comfortably within 16G" is about its own code. ComfyUI's template loads the BF16 files.
- The NVFP4 Turbo file is for Blackwell cards only.
Questions
How much VRAM does Z-Image need?
In ComfyUI's own workflow for Turbo, the largest stage is the diffusion stage at 12.7 GB, and the whole workflow downloads 20.7 GB. Working memory for generation comes on top and grows with resolution. Tongyi-MAI itself says: "fits comfortably within 16G VRAM consumer devices" (Z-Image-Turbo, the vendor's own inference code).
Can I run Z-Image on 8 GB of VRAM?
Yes, with the INT8 ConvRot file: its largest stage, the diffusion stage, is 6.5 GB, which loads on 8 GB with 1.5 GB left for working memory. That is little room to generate at full resolution.
Can I run Z-Image on 12 GB of VRAM?
Yes, with the INT8 ConvRot file: its largest stage, the text encoder, is 8.0 GB, which loads on 12 GB with 5.5 GB left for working memory.
Can I run Z-Image on 16 GB of VRAM?
Yes, with the BF16 file: its largest stage, the diffusion stage, is 12.7 GB, which loads on 16 GB with 3.3 GB left for working memory.
Which Z-Image file should I download?
The highest precision that loads fully on your card. On a 24 GB card that is the BF16 file (12.3 GB). ComfyUI's template downloads BF16 with the BF16 text encoder. FP4 files (NVFP4) are made for Blackwell cards (RTX 50, RTX PRO 6000, DGX Spark); on older cards, use FP8 or INT8.
- Files and sizes: Comfy-Org/z_image_turbo on Hugging Face, read from the Hugging Face API. Template files: ComfyUI's Z-Image guide.
- Model and licence: Tongyi-MAI/Z-Image-Turbo and the Apache 2.0.
- Card memory: manufacturers' specifications, as on each GPU page. The checker covers NVIDIA cards; ComfyUI also runs on AMD and Apple silicon, where file formats and speed differ and we have not sized them.
- Fit rule: a stage loads fully under 85% of usable memory and is tight up to 95%, the same margins the LLM pages use.