Z-Image VRAM requirements in ComfyUI, file by file

Z-Image VRAM in ComfyUI: every file at every precision, which one loads on 12 to 32 GB cards, and what Tongyi-MAI says it needs.

Updated · estimates are labelled as estimates
On this page

Z-Image is Tongyi-MAI's open image model. ComfyUI's own workflow for Turbo downloads 20.7 GB and needs 12.7 GB on the card at its largest stage, the diffusion stage. The smallest card listed here that loads that set with room to spare is the RTX 5060 Ti. With a smaller file, the INT8 ConvRot version, it loads on an RTX 4070.

What Z-Image is

Tongyi-MAI's 6-billion-parameter image model, one of the few here under Apache 2.0. Turbo (November 2025) generates in 8 steps; Base (January 2026) is the undistilled model for fine-tuning and longer sampling.

Turbo: 6.2B · Base: 6.2B Text encoder: Qwen3 4B Apache 2.0 November 2025

Which Z-Image file for your card

For each memory size, the highest-precision diffusion file that loads with room to spare (a file that only just fits is used when nothing else does), and what is left for working memory while it samples. "Streams" means even the smallest file is larger than the card: ComfyUI still runs it by streaming weights from system RAM, more slowly. FP4 files are offered only to the Blackwell cards that run them natively.

Turbo

CardUsableFile that loadsLargest stageLeft for the workVerdict
RTX 4070 12 GB INT8 ConvRot 8.0 GB 5.5 GB Loads fully
RTX 5060 Ti 16 GB BF16 12.7 GB 3.3 GB Loads fully
RTX 4090 24 GB BF16 12.7 GB 11.3 GB Loads fully
RTX 5090 32 GB BF16 12.7 GB 19.3 GB Loads fully
RTX 6000 Ada 48 GB BF16 12.7 GB 35.3 GB Loads fully
H100 SXM 80 GB BF16 12.7 GB 67.3 GB Loads fully
RTX PRO 6000 Blackwell 96 GB BF16 12.7 GB 83.3 GB Loads fully
DGX Spark 126 GB BF16 12.7 GB 113.3 GB Loads fully

ComfyUI's template for Turbo: BF16 diffusion model, BF16 text encoder. Stages: text encoder 8.0 GB, diffusion 12.7 GB. Download 20.7 GB.

Base

CardUsableFile that loadsLargest stageLeft for the workVerdict
RTX 4070 12 GB INT8 ConvRot 8.0 GB 5.5 GB Loads fully
RTX 5060 Ti 16 GB BF16 12.7 GB 3.3 GB Loads fully
RTX 4090 24 GB BF16 12.7 GB 11.3 GB Loads fully
RTX 5090 32 GB BF16 12.7 GB 19.3 GB Loads fully
RTX 6000 Ada 48 GB BF16 12.7 GB 35.3 GB Loads fully
H100 SXM 80 GB BF16 12.7 GB 67.3 GB Loads fully
RTX PRO 6000 Blackwell 96 GB BF16 12.7 GB 83.3 GB Loads fully
DGX Spark 126 GB BF16 12.7 GB 113.3 GB Loads fully

ComfyUI's template for Base: BF16 diffusion model, BF16 text encoder. Stages: text encoder 8.0 GB, diffusion 12.7 GB. Download 20.7 GB.

Check another card, or pick a file by hand, in the ComfyUI VRAM checker.

Every Z-Image file ComfyUI loads

Sizes are the exact byte counts Hugging Face reports. Files marked "template" are the ones ComfyUI's official workflow downloads.

PartPrecisionSizeVersionFile
Diffusion model BF16 · template 12.31 GB Turbo z_image_turbo_bf16.safetensors
Diffusion model INT8 ConvRot 6.20 GB Turbo z_image_turbo_int8_convrot.safetensors
Diffusion model NVFP4 4.51 GB Turbo z_image_turbo_nvfp4.safetensors
Diffusion model BF16 12.31 GB Base z_image_bf16.safetensors
Diffusion model INT8 ConvRot 6.20 GB Base z_image_int8_convrot.safetensors
Text encoder BF16 · template 8.04 GB all qwen_3_4b.safetensors
Text encoder FP8 mixed 5.63 GB all qwen_3_4b_fp8_mixed.safetensors
Text encoder FP4 mixed 3.48 GB all qwen_3_4b_fp4_mixed.safetensors
VAE FP32 · template 0.34 GB all ae.safetensors

What Tongyi-MAI says it needs

  • "fits comfortably within 16G VRAM consumer devices" (Z-Image-Turbo, the vendor's own inference code, source)

Those figures and this page measure different things. A vendor's figure covers its own script at its default resolution, working memory included. The tables here count the weights each stage holds, which is exact, and leave the working memory as the space that remains, because it depends on your resolution and attention kernel and nobody publishes it per model. If generation runs out of memory with a file that loads, lower the resolution before stepping down a precision. ComfyUI's README says it "can run even the biggest open source models on as low as 4GB vram + 8GB ram" by streaming weights (source), so "streams" means slower, not impossible.

LoRA training for Z-Image

No trainer we checked publishes a VRAM figure for training Z-Image yet (kohya's sd-scripts and musubi-tuner, ostris' ai-toolkit, diffusion-pipe, and Tongyi-MAI's own repo). When one does, it goes here.

What Z-Image is good at, and what to watch

Good at

  • Fast drafts: Turbo needs 8 steps where Base takes 28 to 50.
  • Commercial use without conditions beyond the Apache notice.
  • Images up to 2048 by 2048 pixels (Base).

Watch out for

  • Tongyi-MAI's "fits comfortably within 16G" is about its own code. ComfyUI's template loads the BF16 files.
  • The NVFP4 Turbo file is for Blackwell cards only.

Questions

How much VRAM does Z-Image need?

In ComfyUI's own workflow for Turbo, the largest stage is the diffusion stage at 12.7 GB, and the whole workflow downloads 20.7 GB. Working memory for generation comes on top and grows with resolution. Tongyi-MAI itself says: "fits comfortably within 16G VRAM consumer devices" (Z-Image-Turbo, the vendor's own inference code).

Can I run Z-Image on 8 GB of VRAM?

Yes, with the INT8 ConvRot file: its largest stage, the diffusion stage, is 6.5 GB, which loads on 8 GB with 1.5 GB left for working memory. That is little room to generate at full resolution.

Can I run Z-Image on 12 GB of VRAM?

Yes, with the INT8 ConvRot file: its largest stage, the text encoder, is 8.0 GB, which loads on 12 GB with 5.5 GB left for working memory.

Can I run Z-Image on 16 GB of VRAM?

Yes, with the BF16 file: its largest stage, the diffusion stage, is 12.7 GB, which loads on 16 GB with 3.3 GB left for working memory.

Which Z-Image file should I download?

The highest precision that loads fully on your card. On a 24 GB card that is the BF16 file (12.3 GB). ComfyUI's template downloads BF16 with the BF16 text encoder. FP4 files (NVFP4) are made for Blackwell cards (RTX 50, RTX PRO 6000, DGX Spark); on older cards, use FP8 or INT8.

  • Files and sizes: Comfy-Org/z_image_turbo on Hugging Face, read from the Hugging Face API. Template files: ComfyUI's Z-Image guide.
  • Model and licence: Tongyi-MAI/Z-Image-Turbo and the Apache 2.0.
  • Card memory: manufacturers' specifications, as on each GPU page. The checker covers NVIDIA cards; ComfyUI also runs on AMD and Apple silicon, where file formats and speed differ and we have not sized them.
  • Fit rule: a stage loads fully under 85% of usable memory and is tight up to 95%, the same margins the LLM pages use.