Wan 2.2 is Wan's open video model. ComfyUI's own workflow for Text to video, A14B downloads 35.6 GB and needs 14.5 GB on the card at its largest stage, the diffusion stage. The smallest card listed here that loads that set with room to spare is the RTX 4090. With a smaller file, the FP8 scaled version, it loads on an RTX 5060 Ti.
What Wan 2.2 is
Alibaba's Wan 2.2, from July 2025, remains the newest Wan with open weights. The A14B models are two 14-billion-parameter experts that take turns, one for the rough layout and one for detail. TI2V-5B is a single smaller model with a more compressed VAE, and Animate 2 (August 2026) is the character-animation model in the same family.
Text to video, A14B: 14B · Image to video, A14B: 14B · Text and image to video, 5B: 5B · Animate 2: 14B Text encoder: UMT5-XXL Apache 2.0 July 2025
Which Wan 2.2 file for your card
For each memory size, the highest-precision diffusion file that loads with room to spare (a file that only just fits is used when nothing else does), and what is left for working memory while it samples. "Streams" means even the smallest file is larger than the card: ComfyUI still runs it by streaming weights from system RAM, more slowly. FP4 files are offered only to the Blackwell cards that run them natively.
Text to video, A14B (two experts, one on the card at a time)
| Card | Usable | File that loads | Largest stage | Left for the work | Verdict |
|---|---|---|---|---|---|
| RTX 4070 | 12 GB | smallest: FP8 scaled | 14.5 GB | – | Streams |
| RTX 5060 Ti | 16 GB | FP8 scaled | 14.5 GB | 1.5 GB | Loads, tight |
| RTX 4090 | 24 GB | FP8 scaled | 14.5 GB | 9.5 GB | Loads fully |
| RTX 5090 | 32 GB | FP8 scaled | 14.5 GB | 17.5 GB | Loads fully |
| RTX 6000 Ada | 48 GB | FP16 | 28.8 GB | 19.2 GB | Loads fully |
| H100 SXM | 80 GB | FP16 | 28.8 GB | 51.2 GB | Loads fully |
| RTX PRO 6000 Blackwell | 96 GB | FP16 | 28.8 GB | 67.2 GB | Loads fully |
| DGX Spark | 126 GB | FP16 | 28.8 GB | 97.2 GB | Loads fully |
ComfyUI's template for Text to video, A14B: FP8 scaled diffusion model, FP8 scaled text encoder. Stages: text encoder 6.7 GB, diffusion 14.5 GB. Download 35.6 GB.
Image to video, A14B (two experts, one on the card at a time)
| Card | Usable | File that loads | Largest stage | Left for the work | Verdict |
|---|---|---|---|---|---|
| RTX 4070 | 12 GB | smallest: FP8 scaled | 14.5 GB | – | Streams |
| RTX 5060 Ti | 16 GB | FP8 scaled | 14.5 GB | 1.5 GB | Loads, tight |
| RTX 4090 | 24 GB | FP8 scaled | 14.5 GB | 9.5 GB | Loads fully |
| RTX 5090 | 32 GB | FP8 scaled | 14.5 GB | 17.5 GB | Loads fully |
| RTX 6000 Ada | 48 GB | FP16 | 28.8 GB | 19.2 GB | Loads fully |
| H100 SXM | 80 GB | FP16 | 28.8 GB | 51.2 GB | Loads fully |
| RTX PRO 6000 Blackwell | 96 GB | FP16 | 28.8 GB | 67.2 GB | Loads fully |
| DGX Spark | 126 GB | FP16 | 28.8 GB | 97.2 GB | Loads fully |
ComfyUI's template for Image to video, A14B: FP16 diffusion model, FP8 scaled text encoder. Stages: text encoder 6.7 GB, diffusion 28.8 GB. Download 64.2 GB.
Text and image to video, 5B
| Card | Usable | File that loads | Largest stage | Left for the work | Verdict |
|---|---|---|---|---|---|
| RTX 4070 | 12 GB | smallest: FP16 | 11.4 GB | – | Streams |
| RTX 5060 Ti | 16 GB | FP16 | 11.4 GB | 4.6 GB | Loads fully |
| RTX 4090 | 24 GB | FP16 | 11.4 GB | 12.6 GB | Loads fully |
| RTX 5090 | 32 GB | FP16 | 11.4 GB | 20.6 GB | Loads fully |
| RTX 6000 Ada | 48 GB | FP16 | 11.4 GB | 36.6 GB | Loads fully |
| H100 SXM | 80 GB | FP16 | 11.4 GB | 68.6 GB | Loads fully |
| RTX PRO 6000 Blackwell | 96 GB | FP16 | 11.4 GB | 84.6 GB | Loads fully |
| DGX Spark | 126 GB | FP16 | 11.4 GB | 114.6 GB | Loads fully |
ComfyUI's template for Text and image to video, 5B: FP16 diffusion model, FP8 scaled text encoder. Stages: text encoder 6.7 GB, diffusion 11.4 GB. Download 18.2 GB.
Animate 2
| Card | Usable | File that loads | Largest stage | Left for the work | Verdict |
|---|---|---|---|---|---|
| RTX 4070 | 12 GB | smallest: INT8 ConvRot | 16.9 GB | – | Streams |
| RTX 5060 Ti | 16 GB | smallest: INT8 ConvRot | 16.9 GB | – | Streams |
| RTX 4090 | 24 GB | INT8 ConvRot | 16.9 GB | 7.1 GB | Loads fully |
| RTX 5090 | 32 GB | INT8 ConvRot | 16.9 GB | 15.1 GB | Loads fully |
| RTX 6000 Ada | 48 GB | BF16 | 33.0 GB | 15.0 GB | Loads fully |
| H100 SXM | 80 GB | BF16 | 33.0 GB | 47.0 GB | Loads fully |
| RTX PRO 6000 Blackwell | 96 GB | BF16 | 33.0 GB | 63.0 GB | Loads fully |
| DGX Spark | 126 GB | BF16 | 33.0 GB | 93.0 GB | Loads fully |
ComfyUI's template for Animate 2: INT8 ConvRot diffusion model, FP8 scaled text encoder. Stages: text encoder 8.0 GB, diffusion 16.9 GB. Download 24.9 GB.
Check another card, or pick a file by hand, in the ComfyUI VRAM checker.
Every Wan 2.2 file ComfyUI loads
Sizes are the exact byte counts Hugging Face reports. Files marked "template" are the ones ComfyUI's official workflow downloads.
| Part | Precision | Size | Version | File |
|---|---|---|---|---|
| Diffusion model | FP16 | 28.58 GB | Text to video, A14B | wan2.2_t2v_high_noise_14B_fp16.safetensors |
| Diffusion model | FP8 scaled · template | 14.29 GB | Text to video, A14B | wan2.2_t2v_high_noise_14B_fp8_scaled.safetensors |
| Diffusion model | FP16 · template | 28.58 GB | Image to video, A14B | wan2.2_i2v_high_noise_14B_fp16.safetensors |
| Diffusion model | FP8 scaled | 14.29 GB | Image to video, A14B | wan2.2_i2v_high_noise_14B_fp8_scaled.safetensors |
| Diffusion model | FP16 · template | 10.00 GB | Text and image to video, 5B | wan2.2_ti2v_5B_fp16.safetensors |
| Diffusion model | BF16 | 32.79 GB | Animate 2 | wan_animate_2_bf16.safetensors |
| Diffusion model | INT8 ConvRot · template | 16.65 GB | Animate 2 | wan_animate_2_int8_convrot.safetensors |
| Text encoder | FP16 | 11.37 GB | all | umt5_xxl_fp16.safetensors |
| Text encoder | FP8 scaled · template | 6.74 GB | all | umt5_xxl_fp8_e4m3fn_scaled.safetensors |
| VAE | BF16 · template | 0.25 GB | Text to video, A14B | wan_2.1_vae.safetensors |
| VAE | BF16 · template | 0.25 GB | Image to video, A14B | wan_2.1_vae.safetensors |
| VAE | FP16 · template | 1.41 GB | Text and image to video, 5B | wan2.2_vae.safetensors |
| VAE | BF16 · template | 0.25 GB | Animate 2 | Wan2_1_VAE_bf16.safetensors |
| Vision encoder | FP16 · template | 1.26 GB | Animate 2 | clip_vision_h.safetensors |
What Wan says it needs
- "This command can run on a GPU with at least 80GB VRAM." (A14B models, Wan's own script, source)
- "This command can run on a GPU with at least 80GB VRAM." (A14B models, Wan's own script, source)
- "This command can run on a GPU with at least 24GB VRAM (e.g, RTX 4090 GPU)." (TI2V-5B at 720p, Wan's own script with offloading, source)
Those figures and this page measure different things. A vendor's figure covers its own script at its default resolution and length, working memory included. The tables here count the weights each stage holds, which is exact, and leave the working memory as the space that remains, because it depends on your resolution, frame count and attention kernel and nobody publishes it per model. If generation runs out of memory with a file that loads, lower the resolution or the frame count before stepping down a precision. ComfyUI's README says it "can run even the biggest open source models on as low as 4GB vram + 8GB ram" by streaming weights (source), so "streams" means slower, not impossible.
Wan 2.5, 2.6, 2.7 and 3.0 have no open weights (checked 2026-10-02). None has open weights on Hugging Face (Wan-AI), ModelScope or GitHub (Wan-Video). Wan offers them through Alibaba Cloud Model Studio, the Wan website, the Qwen app and Qwen Cloud, and has not said they will be released.
LoRA training for Wan 2.2
Only configurations a trainer's own documentation gives a VRAM figure for. Settings change the figure a great deal, so read the setting column before trusting the number.
| Trainer | VRAM | Setting | Source |
|---|---|---|---|
| musubi-tuner · Text to video, A14B | 24 GB | 720×1280 images, FP8 model, block swapping | docs |
| ai-toolkit · Text to video, A14B | 24 GB | images at 512 to 1024 pixels, 4-bit model; video is "very VRAM intensive" | docs |
What Wan 2.2 is good at, and what to watch
Good at
- Apache 2.0, so commercial use is open.
- 480p and 720p clips of about five seconds.
- LoRA training on 24 GB cards, documented by musubi-tuner and ai-toolkit.
Watch out for
- Two experts mean two downloads of the same size, even though only one is on the card at a time.
- Wan's own script for the A14B models asks for an 80 GB GPU; ComfyUI gets by with far less by moving models in and out.
- Wan 2.5, 2.6, 2.7 and 3.0 are hosted only, with no open weights.
Questions
How much VRAM does Wan 2.2 need?
In ComfyUI's own workflow for Text to video, A14B, the largest stage is the diffusion stage at 14.5 GB, and the whole workflow downloads 35.6 GB. Working memory for generation comes on top and grows with resolution and length. Wan itself says: "This command can run on a GPU with at least 80GB VRAM." (A14B models, Wan's own script).
Can I run Wan 2.2 on 8 GB of VRAM?
Not fully. The smallest file here, FP8 scaled at 14.3 GB, needs 14.5 GB at its largest stage, more than 8 GB holds. ComfyUI will still run it by streaming weights from system RAM, which works but is slower.
Can I run Wan 2.2 on 12 GB of VRAM?
Not fully. The smallest file here, FP8 scaled at 14.3 GB, needs 14.5 GB at its largest stage, more than 12 GB holds. ComfyUI will still run it by streaming weights from system RAM, which works but is slower.
Can I run Wan 2.2 on 16 GB of VRAM?
Yes, with the FP8 scaled file: its largest stage, the diffusion stage, is 14.5 GB, which loads on 16 GB with 1.5 GB left for working memory, a tight fit. That is little room to generate at full resolution.
Which Wan 2.2 file should I download?
The highest precision that loads fully on your card. On a 24 GB card that is the FP8 scaled file (14.3 GB). ComfyUI's template downloads FP8 scaled with the FP8 scaled text encoder. FP4 files (NVFP4) are made for Blackwell cards (RTX 50, RTX PRO 6000, DGX Spark); on older cards, use FP8 or INT8.
Can I run Wan 2.5, 2.6, 2.7 or 3.0 locally?
No. None of them has open weights: they are not on Hugging Face, ModelScope or Wan's GitHub, and Wan offers them only through Alibaba Cloud Model Studio, the Wan website, the Qwen app and Qwen Cloud. Wan 2.2, including Animate 2, is the newest you can download.
- Files and sizes: Comfy-Org/Wan_2.2_ComfyUI_Repackaged on Hugging Face, read from the Hugging Face API. Template files: ComfyUI's Wan 2.2 guide.
- Model and licence: Wan-AI/Wan2.2-T2V-A14B and the Apache 2.0.
- Card memory: manufacturers' specifications, as on each GPU page. The checker covers NVIDIA cards; ComfyUI also runs on AMD and Apple silicon, where file formats and speed differ and we have not sized them.
- Fit rule: a stage loads fully under 85% of usable memory and is tight up to 95%, the same margins the LLM pages use.