MiniMax H3 is MiniMax's open video model with sound. ComfyUI's own workflow downloads 40.1 GB and needs 24.4 GB on the card at its largest stage, the diffusion stage. The smallest card listed here that loads that set with room to spare is the RTX 5090. With a smaller file, the W6A8, pruned version, it loads on an RTX 4090.
The licence does not apply in the United States, the European Union, the United Kingdom and South Korea. If you are there, this licence does not cover you.
What MiniMax H3 is
MiniMax's open video model from August 2026, and the only one here that makes the soundtrack with the picture: stereo audio, 4 to 15 seconds at 24 frames a second. Its 33-billion-parameter transformer is paired with a 32-billion-parameter vision-language model as its text encoder, so even the encoder is the size of a large language model.
33B parameters Text encoder: Qwen3-VL 32B MiniMax H3 Community License August 2026
Before you download
Where the licence applies, commercial use is allowed; products above US$20 million a year need MiniMax's written authorisation, and commercial interfaces must show "MiniMax H3". The licence does not apply at all in the United States, the European Union, the United Kingdom and South Korea. Read the MiniMax H3 Community License itself before using it for work.
Which MiniMax H3 file for your card
For each memory size, the highest-precision diffusion file that loads with room to spare (a file that only just fits is used when nothing else does), and what is left for working memory while it samples. "Streams" means even the smallest file is larger than the card: ComfyUI still runs it by streaming weights from system RAM, more slowly. FP4 files are offered only to the Blackwell cards that run them natively.
| Card | Usable | File that loads | Largest stage | Left for the work | Verdict |
|---|---|---|---|---|---|
| RTX 4070 | 12 GB | smallest: W6A8, pruned | 19.4 GB | – | Streams |
| RTX 5060 Ti | 16 GB | smallest: W6A8, pruned | 19.4 GB | – | Streams |
| RTX 4090 | 24 GB | W6A8, pruned | 19.4 GB | 4.6 GB | Loads fully |
| RTX 5090 | 32 GB | FP8 scaled, pruned | 27.1 GB | 7.6 GB | Loads fully |
| RTX 6000 Ada | 48 GB | FP8 scaled, pruned | 27.1 GB | 23.6 GB | Loads fully |
| H100 SXM | 80 GB | BF16, pruned | 51.5 GB | 36.3 GB | Loads fully |
| RTX PRO 6000 Blackwell | 96 GB | BF16, pruned | 51.5 GB | 52.3 GB | Loads fully |
| DGX Spark | 126 GB | BF16, pruned | 51.5 GB | 82.3 GB | Loads fully |
ComfyUI's template: INT8 ConvRot, pruned diffusion model, NVFP4 AWQ text encoder. Stages: text encoder 15.7 GB, diffusion 24.4 GB. Download 40.1 GB.
Check another card, or pick a file by hand, in the ComfyUI VRAM checker.
Every MiniMax H3 file ComfyUI loads
Sizes are the exact byte counts Hugging Face reports. Files marked "template" are the ones ComfyUI's official workflow downloads.
| Part | Precision | Size | File |
|---|---|---|---|
| Diffusion model | BF16 | 66.28 GB | minimax_h3_fl2va_bf16.safetensors |
| Diffusion model | INT8 ConvRot | 34.04 GB | minimax_h3_fl2va_int8_convrot.safetensors |
| Diffusion model | BF16, pruned | 40.23 GB | minimax_h3_fl2va_pruned_bf16.safetensors |
| Diffusion model | FP8 scaled, pruned | 20.96 GB | minimax_h3_fl2va_pruned_fp8_scaled.safetensors |
| Diffusion model | INT8 ConvRot, pruned · template | 20.97 GB | minimax_h3_fl2va_pruned_int8_convrot.safetensors |
| Diffusion model | W6A8, pruned | 15.98 GB | minimax_h3_fl2va_pruned_w6a8.safetensors |
| Text encoder | BF16 | 51.51 GB | qwen3vl_32b_minimax_h3_bf16.safetensors |
| Text encoder | INT8 ConvRot | 27.14 GB | qwen3vl_32b_minimax_h3_int8_convrot.safetensors |
| Text encoder | NVFP4 AWQ · template | 15.69 GB | qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors |
| VAE | INT8 ConvRot · template | 2.81 GB | minimax_h3_video_vae_int8_convrot.safetensors |
| Audio VAE | FP32 · template | 0.61 GB | minimax_h3_audio_vae_fp32.safetensors |
What MiniMax says it needs
MiniMax states no VRAM figure. Its only hardware hint is a serving example on four GPUs. The closest it comes: "sglang serve --model-path MiniMaxAI/MiniMax-H3 --num-gpus 4 --ulysses-degree 4" (source).
Those figures and this page measure different things. A vendor's figure covers its own script at its default resolution and length, working memory included. The tables here count the weights each stage holds, which is exact, and leave the working memory as the space that remains, because it depends on your resolution, frame count and attention kernel and nobody publishes it per model. If generation runs out of memory with a file that loads, lower the resolution or the frame count before stepping down a precision. ComfyUI's README says it "can run even the biggest open source models on as low as 4GB vram + 8GB ram" by streaming weights (source), so "streams" means slower, not impossible.
LoRA training for MiniMax H3
Only configurations a trainer's own documentation gives a VRAM figure for. Settings change the figure a great deal, so read the setting column before trusting the number.
| Trainer | VRAM | Setting | Source |
|---|---|---|---|
| diffusion-pipe | 24 GB | 768-pixel default, the pruned INT8 ConvRot model, block swapping | docs |
What MiniMax H3 is good at, and what to watch
Good at
- Video with synchronised sound in one pass, from text, from first and last frames, or from reference images.
- Comfy-Org's pruned files, which drop the parameters the model can rebuild and cut the transformer by more than a third.
- A Turbo LoRA from lightx2v that ComfyUI's template uses for fewer sampling steps.
Watch out for
- The licence does not apply in the United States, the European Union, the United Kingdom or South Korea.
- Comfy's docs describe a separate licence for local commercial use, sold through Comfy. Read both before using it for work.
- Comfy-Org recommends its INT8 ConvRot files where your PyTorch is built for CUDA 13.0, and FP8 otherwise.
Questions
How much VRAM does MiniMax H3 need?
In ComfyUI's own workflow, the largest stage is the diffusion stage at 24.4 GB, and the whole workflow downloads 40.1 GB. Working memory for generation comes on top and grows with resolution and length. MiniMax publishes no VRAM figure.
Can I run MiniMax H3 on 8 GB of VRAM?
Not fully. The smallest file here, W6A8, pruned at 16.0 GB, needs 19.4 GB at its largest stage, more than 8 GB holds. ComfyUI will still run it by streaming weights from system RAM, which works but is slower.
Can I run MiniMax H3 on 12 GB of VRAM?
Not fully. The smallest file here, W6A8, pruned at 16.0 GB, needs 19.4 GB at its largest stage, more than 12 GB holds. ComfyUI will still run it by streaming weights from system RAM, which works but is slower.
Can I run MiniMax H3 on 16 GB of VRAM?
Not fully. The smallest file here, W6A8, pruned at 16.0 GB, needs 19.4 GB at its largest stage, more than 16 GB holds. ComfyUI will still run it by streaming weights from system RAM, which works but is slower.
Which MiniMax H3 file should I download?
The highest precision that loads fully on your card. On a 24 GB card that is the W6A8, pruned file (16.0 GB). ComfyUI's template downloads INT8 ConvRot, pruned with the NVFP4 AWQ text encoder. FP4 files (NVFP4) are made for Blackwell cards (RTX 50, RTX PRO 6000, DGX Spark); on older cards, use FP8 or INT8.
Can I use MiniMax H3 in the US or the EU?
Not under its open licence. The MiniMax H3 Community License lists the United States, the European Union, the United Kingdom and South Korea as excluded territories, where it does not apply. Elsewhere, commercial use is allowed, and products above US$20 million a year need MiniMax's written authorisation.
- Files and sizes: Comfy-Org/MiniMax-H3 on Hugging Face, read from the Hugging Face API. Template files: ComfyUI's MiniMax H3 guide.
- Model and licence: MiniMaxAI/MiniMax-H3 and the MiniMax H3 Community License.
- Card memory: manufacturers' specifications, as on each GPU page. The checker covers NVIDIA cards; ComfyUI also runs on AMD and Apple silicon, where file formats and speed differ and we have not sized them.
- Fit rule: a stage loads fully under 85% of usable memory and is tight up to 95%, the same margins the LLM pages use.