MINIMAX H3 / COMFYUI / INDEPENDENT GUIDE

MiniMax H3 ComfyUI Workflow

Download the official model files, place them correctly, and make the first run with a practical workflow, settings, troubleshooting, and prompt guide.

Last verified Aug 7, 2026Not affiliated with MiniMax or ComfyUI
first-run checklist

$ comfyui --check-h3

01 download the four model files

02 place files in matching folders

03 update ComfyUI

04 open the official T2V template

05 start at 1344 × 768 / 5s

01Download modelsUse the four official files
02Place modelsKeep each file in its folder
03Update ComfyUIUse a current build
04Open workflowStart with official T2V
05Run baseline1344 × 768 / 5 seconds

Use this page as a field guide. The model, workflow, and directory facts below are linked to official sources. Hardware advice is a starting point, not a cross-GPU benchmark.

01 / Workflows

Choose the first run

Start with T2V. Move to image or frame references once the local install is clean.

A / DEFAULTT2V

Text to Video

Describe the subject, evolving action, camera, and audio in one structured prompt.

  • Input: text prompt
  • Best first check: official baseline
  • Template: Comfy-Org T2V
B / REFERENCEI2V

Image to Video

Use an image as the first frame and describe what changes after the reference frame.

  • Input: first-frame image
  • Prompt: preserve then animate
  • Guide: reference alignment
C / TRANSITIONFL2V

First + Last Frame

Describe an observable transition from the first frame toward the final frame.

  • Input: first and final frame
  • Prompt: transition path
  • Guide: H3 prompt structure
02 / Download & Model Setup

Download the files

Download the official model files, then place each one in the matching ComfyUI folder. The current official T2V and I2V template manifests list all four files.

ComfyUI/
└── models/
    ├── vae/
    │   ├── minimax_h3_video_vae_fp16.safetensors
    │   └── minimax_h3_audio_vae_fp32.safetensors
    ├── diffusion_models/
    │   └── minimax_h3_fl2va_pruned_int8_convrot.safetensors
    └── text_encoders/
        └── qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors
Video VAEminimax_h3_video_vae_fp16.safetensors
RequiredT2V + I2V
Foldermodels/vaeUseRequired for the current official T2V and I2V templates.Download file
Audio VAEminimax_h3_audio_vae_fp32.safetensors
RequiredT2V + I2V
Foldermodels/vaeUseRequired for the current official T2V and I2V templates.Download file
H3 diffusion modelminimax_h3_fl2va_pruned_int8_convrot.safetensors
RequiredT2V + I2V
Foldermodels/diffusion_modelsUseRequired for the current official T2V and I2V templates.Download file
Qwen text encoderqwen3vl_32b_minimax_h3_nvfp4_awq.safetensors
RequiredT2V + I2V
Foldermodels/text_encodersUseRequired for the current official T2V and I2V templates.Download file
Download checklist4 files · T2V + I2V
  1. 01Download all four files from the official model links above.
  2. 02Copy each file into the exact ComfyUI/models/ subfolder shown.
  3. 03Keep every filename unchanged; restart or rescan ComfyUI after copying.
  4. 04Open the official T2V or I2V template only after all four files are visible.

These download URLs point to the Comfy-Org MiniMax H3 repository on Hugging Face. After the files are visible, open the official T2V template or official I2V template . If nodes or model selectors are missing, update ComfyUI before changing random settings.

03 / Best settings

Start from a known baseline

The official template gives a clean starting point. The other two cards are editorial starting points, not benchmarks.

Official baseline1344 × 768 / 5s

20 steps · res_multistep sampler · simple scheduler. Use this as the first comparison point before tuning for your hardware.

StarterRecommended
736 × 416 / 5s

A smaller official-size starting point for a first local check.

Recommended starting point
BalancedRecommended
1056 × 608 / 5s

A middle size from the official resolution reference before the full baseline.

Recommended starting point
OfficialSource
1344 × 768 / 5s

The current T2V template baseline with 20 steps, res_multistep, and simple.

Official baseline
VRAM / OOM

There is no universal minimum VRAM promise here. If you hit OOM, lower resolution and duration first, then check ComfyUI version, attention/offload, and model precision.

Dynamic VRAM / SageAttention

Treat these as compatibility paths, not magic switches. Update ComfyUI, confirm the workflow loads, then test one attention or offload change at a time.

04 / Troubleshooting

Fix the first failure

Classify the error before changing settings. One clean run is more useful than ten random tweaks.

Missing model

Model selector is empty

  1. Check the exact filename.
  2. Move it into the matching folder.
  3. Restart or rescan ComfyUI.
Missing node

Workflow cannot load

  1. Update ComfyUI.
  2. Open the current official template.
  3. Compare the node names again.
OOM

Run stops in sampling

  1. Lower resolution and duration.
  2. Check attention/offload options.
  3. Change one setting, then retry.

For version and workflow-template issues, start with the ComfyUI update guide . For model and node bugs, use the official support record .

05 / Prompt builder

Build a structured H3 prompt

This formatter stays in your browser. It uses curated director variations instead of random word fragments.

Local only. No API, no upload, no saved prompt.

6 presets × 3 variants
Preview / variant AT2V
integrated_multimodal_description: Cinematic realism. medium shot of a woman in a red raincoat in a neon-lit city street after dark. a woman in a red raincoat walks through the rain and turns toward the camera. The camera uses slow push in, keeping the motion clear and coherent. Neon night.  End on the woman pauses beneath a glowing sign. no text, no logos, keep the face consistent
overall_soundscape: rainfall, footsteps on wet pavement, distant traffic
non_diegetic_music: N/A
Browser only

Prompt syntax is based on the H3 Base Prompt Guide . Full Reference / Ref2VA uses a separate Reference Prompt Guide and is not simplified here.

06 / Static examples

Read the shape

These examples remain in the HTML so the guide is useful even before the builder loads.

Cinematic T2V / rain and neonCopyable text
integrated_multimodal_description: Cinematic realism. Medium shot of a woman in a red raincoat in a neon-lit city street after dark. She walks through the rain and turns toward the camera. The camera uses a slow push in, keeping the motion clear and coherent. Neon night. End on the woman pausing beneath a glowing sign. no text, no logos, keep the face consistent.
overall_soundscape: rainfall, footsteps on wet pavement, distant traffic
non_diegetic_music: N/A
I2V / first-frame continuityCopyable text
<Picture 1> is the exact first frame at 0.00 seconds. Preserve the subject, composition, clothing, and environment from the reference; describe only what changes after the first frame.

integrated_multimodal_description: A close-up of the ceramic cup remains on the wooden table as steam slowly rises and the camera performs a restrained push in. Warm window light stays consistent.
overall_soundscape: quiet room tone, soft kettle hiss
non_diegetic_music: N/A
Dialogue / bound speakerCopyable text
integrated_multimodal_description: Cinematic realism, medium shot of a woman at a train platform at dusk. The woman turns toward the arriving train and says: <d>[English] I get off at the next station.</d> Keep the line associated with the named speaker and keep the camera on her readable facial performance.
overall_soundscape: station ambience, rail vibration, footsteps
non_diegetic_music: N/A
07 / FAQ

Questions after the first run

Short answers for the queries that usually appear when a new workflow meets an old install.

How do I run MiniMax H3 in ComfyUI?

Update ComfyUI, place the official model files in the matching folders, open the official H3 workflow template, and start with the 1344 × 768 / 5s baseline.

Where should MiniMax H3 model files go?

VAE files go in models/vae, the diffusion model goes in models/diffusion_models, and the text encoder goes in models/text_encoders. Use the exact filenames shown in the Models section.

How much VRAM does MiniMax H3 need?

This guide does not publish a universal minimum. Hardware, precision, resolution, duration, and offload choices change the result. For OOM, lower resolution and duration first, then check the install and attention path.

What is the best starting setting?

Use the current official T2V baseline: 1344 × 768, 5 seconds, 20 steps, res_multistep sampler, and simple scheduler. Starter and Balanced cards are editorial starting points for local testing.

Does the Prompt Builder send my prompt anywhere?

No. The builder runs in the browser, does not call an AI API, does not save text, and does not send your free-text fields to analytics.

Does this page support Full Reference / Ref2VA?

The page links the official Reference Prompt Guide, but the first builder supports T2V, I2V, and First + Last Frame only. It does not pretend to format the stricter multi-reference syntax.

08 / Sources

Check the source

This page is an independent editorial guide. Verify fast-moving details against the official pages before a production run.

MiniMax and ComfyUI are trademarks and projects of their respective owners. This site is an independent guide and is not affiliated with either project.