FLUX.2 is Black Forest Labs' open image model. ComfyUI's own workflow for klein 4B downloads 12.5 GB and needs 8.0 GB on the card at its largest stage, the text encoder. The smallest card listed here that loads that set with room to spare is the RTX 4070.
What FLUX.2 is
Black Forest Labs' second FLUX generation: dev (November 2025), a 32-billion-parameter model, and klein (January 2026), 4B and 9B models built for consumer cards. Only klein 4B is Apache 2.0; klein 9B and dev are for non-commercial use.
klein 4B: 4B · klein 9B: 9B · dev: 32B Text encoder: Qwen3 (klein) or Mistral Small (dev) FLUX Non-Commercial License January 2026
Before you download
The weights are for non-commercial use; commercial use needs a paid licence from Black Forest Labs. Read the FLUX Non-Commercial License itself before using it for work. klein 4B is under Apache 2.0 instead: the weights can be used commercially.
Which FLUX.2 file for your card
For each memory size, the highest-precision diffusion file that loads with room to spare (a file that only just fits is used when nothing else does), and what is left for working memory while it samples. "Streams" means even the smallest file is larger than the card: ComfyUI still runs it by streaming weights from system RAM, more slowly. FP4 files are offered only to the Blackwell cards that run them natively.
klein 4B
| Card | Usable | File that loads | Largest stage | Left for the work | Verdict |
|---|---|---|---|---|---|
| RTX 4070 | 12 GB | BF16 | 8.1 GB | 3.9 GB | Loads fully |
| RTX 5060 Ti | 16 GB | BF16 | 8.1 GB | 7.9 GB | Loads fully |
| RTX 4090 | 24 GB | BF16 | 8.1 GB | 15.9 GB | Loads fully |
| RTX 5090 | 32 GB | BF16 | 8.1 GB | 23.9 GB | Loads fully |
| RTX 6000 Ada | 48 GB | BF16 | 8.1 GB | 39.9 GB | Loads fully |
| H100 SXM | 80 GB | BF16 | 8.1 GB | 71.9 GB | Loads fully |
| RTX PRO 6000 Blackwell | 96 GB | BF16 | 8.1 GB | 87.9 GB | Loads fully |
| DGX Spark | 126 GB | BF16 | 8.1 GB | 117.9 GB | Loads fully |
ComfyUI's template for klein 4B: FP8 diffusion model, BF16 text encoder. Stages: text encoder 8.0 GB, diffusion 4.4 GB. Download 12.5 GB.
klein 9B
| Card | Usable | File that loads | Largest stage | Left for the work | Verdict |
|---|---|---|---|---|---|
| RTX 4070 | 12 GB | FP8 | 9.8 GB | 2.2 GB | Loads fully |
| RTX 5060 Ti | 16 GB | FP8 | 9.8 GB | 6.2 GB | Loads fully |
| RTX 4090 | 24 GB | BF16 | 18.5 GB | 5.5 GB | Loads fully |
| RTX 5090 | 32 GB | BF16 | 18.5 GB | 13.5 GB | Loads fully |
| RTX 6000 Ada | 48 GB | BF16 | 18.5 GB | 29.5 GB | Loads fully |
| H100 SXM | 80 GB | BF16 | 18.5 GB | 61.5 GB | Loads fully |
| RTX PRO 6000 Blackwell | 96 GB | BF16 | 18.5 GB | 77.5 GB | Loads fully |
| DGX Spark | 126 GB | BF16 | 18.5 GB | 107.5 GB | Loads fully |
ComfyUI's template for klein 9B: FP8 diffusion model, FP8 mixed text encoder. Stages: text encoder 8.7 GB, diffusion 9.8 GB. Download 18.4 GB.
dev
| Card | Usable | File that loads | Largest stage | Left for the work | Verdict |
|---|---|---|---|---|---|
| RTX 4070 | 12 GB | smallest: FP8 mixed | 35.8 GB | – | Streams |
| RTX 5060 Ti | 16 GB | smallest: NVFP4 | 21.4 GB | – | Streams |
| RTX 4090 | 24 GB | smallest: FP8 mixed | 35.8 GB | – | Streams |
| RTX 5090 | 32 GB | NVFP4 mixed | 23.1 GB | 8.9 GB | Loads fully |
| RTX 6000 Ada | 48 GB | FP8 mixed | 35.8 GB | 12.2 GB | Loads fully |
| H100 SXM | 80 GB | BF16 | 64.8 GB | 15.2 GB | Loads fully |
| RTX PRO 6000 Blackwell | 96 GB | BF16 | 64.8 GB | 31.2 GB | Loads fully |
| DGX Spark | 126 GB | BF16 | 64.8 GB | 61.2 GB | Loads fully |
ComfyUI's template for dev: FP8 mixed diffusion model, BF16 text encoder. Stages: text encoder 35.6 GB, diffusion 35.8 GB. Download 71.4 GB.
Check another card, or pick a file by hand, in the ComfyUI VRAM checker.
Every FLUX.2 file ComfyUI loads
Sizes are the exact byte counts Hugging Face reports. Files marked "template" are the ones ComfyUI's official workflow downloads.
| Part | Precision | Size | Version | File |
|---|---|---|---|---|
| Diffusion model | BF16 | 7.75 GB | klein 4B | flux-2-klein-4b.safetensors |
| Diffusion model | FP8 · template | 4.07 GB | klein 4B | flux-2-klein-4b-fp8.safetensors |
| Diffusion model | BF16 | 18.16 GB | klein 9B | flux-2-klein-9b.safetensors |
| Diffusion model | FP8 · template | 9.43 GB | klein 9B | flux-2-klein-9b-fp8.safetensors |
| Diffusion model | NVFP4 | 5.76 GB | klein 9B | flux-2-klein-9b-nvfp4.safetensors |
| Diffusion model | BF16 | 64.45 GB | dev | flux2-dev.safetensors |
| Diffusion model | FP8 mixed · template | 35.46 GB | dev | flux2_dev_fp8mixed.safetensors |
| Diffusion model | NVFP4 mixed | 22.77 GB | dev | flux2-dev-nvfp4-mixed.safetensors |
| Diffusion model | NVFP4 | 21.04 GB | dev | flux2-dev-nvfp4.safetensors |
| Text encoder | BF16 · template | 8.04 GB | klein 4B | qwen_3_4b.safetensors |
| Text encoder | FP4 | 3.85 GB | klein 4B | qwen_3_4b_fp4_flux2.safetensors |
| Text encoder | BF16 | 16.38 GB | klein 9B | qwen_3_8b.safetensors |
| Text encoder | FP8 mixed · template | 8.66 GB | klein 9B | qwen_3_8b_fp8mixed.safetensors |
| Text encoder | FP4 mixed | 6.80 GB | klein 9B | qwen_3_8b_fp4mixed.safetensors |
| Text encoder | BF16 · template | 35.58 GB | dev | mistral_3_small_flux2_bf16.safetensors |
| Text encoder | FP8 | 18.03 GB | dev | mistral_3_small_flux2_fp8.safetensors |
| Text encoder | FP4 mixed | 12.28 GB | dev | mistral_3_small_flux2_fp4_mixed.safetensors |
| VAE | FP32 · template | 0.34 GB | all | flux2-vae.safetensors |
What Black Forest Labs says it needs
- "Runs on consumer GPUs (~13GB VRAM)." (klein 4B, the model card, source)
- "Klein 4B fits in ~8GB VRAM (RTX 3090/4070 and up)" (klein 4B, the GitHub README (it disagrees with the card), source)
- "The FLUX.2 [klein] 9B model fits in ~29GB VRAM and is accessible on NVIDIA RTX 4090 and above." (klein 9B, the model card (a 4090 has 24 GB, so this assumes offloading), source)
- "Those with 24-32GB of VRAM can use the model with 4-bit quantization" (dev, Black Forest Labs' diffusers guide, source)
Those figures and this page measure different things. A vendor's figure covers its own script at its default resolution, working memory included. The tables here count the weights each stage holds, which is exact, and leave the working memory as the space that remains, because it depends on your resolution and attention kernel and nobody publishes it per model. If generation runs out of memory with a file that loads, lower the resolution before stepping down a precision. ComfyUI's README says it "can run even the biggest open source models on as low as 4GB vram + 8GB ram" by streaming weights (source), so "streams" means slower, not impossible.
LoRA training for FLUX.2
Only configurations a trainer's own documentation gives a VRAM figure for. Settings change the figure a great deal, so read the setting column before trusting the number.
| Trainer | VRAM | Setting | Source |
|---|---|---|---|
| diffusion-pipe · dev | 48 GB | dev, model held in FP8, no block swapping | docs |
What FLUX.2 is good at, and what to watch
Good at
- klein 4B: a small, commercially usable model with FP8 weights from Black Forest Labs.
- dev: images up to 4 megapixels.
- FP8 and NVFP4 files published by Black Forest Labs itself.
Watch out for
- dev's text encoder, Mistral Small 3, is as large in BF16 as the FP8 dev model itself.
- Black Forest Labs' own figures disagree: the klein 4B card says about 13 GB, its GitHub README about 8 GB, and the klein 9B card says about 29 GB while calling an RTX 4090 enough.
- klein 9B and dev need a paid licence for commercial use.
Questions
How much VRAM does FLUX.2 need?
In ComfyUI's own workflow for klein 4B, the largest stage is the text encoder at 8.0 GB, and the whole workflow downloads 12.5 GB. Working memory for generation comes on top and grows with resolution. Black Forest Labs itself says: "Runs on consumer GPUs (~13GB VRAM)." (klein 4B, the model card).
Can I run FLUX.2 on 8 GB of VRAM?
Yes, with the FP8 file: its largest stage, the diffusion stage, is 4.4 GB, which loads on 8 GB with 3.6 GB left for working memory.
Can I run FLUX.2 on 12 GB of VRAM?
Yes, with the BF16 file: its largest stage, the diffusion stage, is 8.1 GB, which loads on 12 GB with 3.9 GB left for working memory.
Can I run FLUX.2 on 16 GB of VRAM?
Yes, with the BF16 file: its largest stage, the diffusion stage, is 8.1 GB, which loads on 16 GB with 7.9 GB left for working memory.
Which FLUX.2 file should I download?
The highest precision that loads fully on your card. On a 24 GB card that is the BF16 file (7.8 GB). ComfyUI's template downloads FP8 with the BF16 text encoder. FP4 files (NVFP4) are made for Blackwell cards (RTX 50, RTX PRO 6000, DGX Spark); on older cards, use FP8 or INT8.
- Files and sizes: Comfy-Org/flux2-dev on Hugging Face, plus the vendor's own ComfyUI-ready files, read from the Hugging Face API. Template files: ComfyUI's FLUX.2 guide.
- Model and licence: black-forest-labs/FLUX.2-klein-4B and the FLUX Non-Commercial License.
- Card memory: manufacturers' specifications, as on each GPU page. The checker covers NVIDIA cards; ComfyUI also runs on AMD and Apple silicon, where file formats and speed differ and we have not sized them.
- Fit rule: a stage loads fully under 85% of usable memory and is tight up to 95%, the same margins the LLM pages use.