Feedback
AI Ad Video Example
Loading...
comfyui minimax h3
Turn a prompt, still, or clip into video with synced stereo sound on the comfyui minimax h3 graph — open weights, local runs, full node control.
All Tools
Discover our comprehensive AI-powered animation toolkit
MiniMax H3
MiniMax H3 AI Video Generator
Seedance 2.5
The Future of AI Video Is Here.

Seedance 2.0
The Future of AI Video Is Here.

Veo3.1
Create Stunning Videos with Veo3.1
MiniMax H3 video generator
MiniMax H3 AI Video Generator

Seedance 2.0
The Future of AI Video Is Here.

Kling Motion Control
Turn reference images into amazing motion videos in minutes
Sora2 AI
Advanced AI Video Generator for High-Quality Videos
Why Creators Pick the comfyui minimax h3 Workflow
Built on open weights, the comfyui minimax h3 workflow brings MiniMax's omni-modal engine into ComfyUI as a node graph you own. A single context absorbs text, pictures, footage, and sound at once, so clips arrive with matching stereo audio — dialogue, effects, and score produced in the same pass. Results land at up to 2K and 24fps for roughly 15 seconds, and every diffusion setting stays in your hands.
- Built-In Stereo SoundSpeech, effects, and music are rendered alongside the picture and muxed into one MP4 — the comfyui minimax h3 graph keeps them in sync from a single pass.
- Open Weights, Zero CapsBecause the comfyui minimax h3 checkpoint runs on your own GPU, you set resolution, length, and every diffusion knob without hitting an API ceiling.
- Mix Any Reference TypeFeed text, stills, footage, and audio into one run and pin down a face, a look, a motion path, a camera move, or a voice through the comfyui minimax h3 nodes.
Set Up the comfyui minimax h3 Workflow in Three Steps
Three quick steps take you from a fresh ComfyUI install to open-weight video with synced audio via the comfyui minimax h3 graph.
Key Strengths of the comfyui minimax h3 Workflow
From three ready-made graphs to open-weight multimodal inference, synced stereo sound, reference-driven direction, and optional Sage Attention acceleration, the comfyui minimax h3 workflow covers a full local production pipeline.
Three Ready-Made Graphs
Text-to-video, image-to-video, and reference-to-video samples ship in the comfyui minimax h3 template library, each one wired for a different generation mode.
One Shared Context
Text, pictures, footage, and sound are all read together by the comfyui minimax h3 model, letting you blend reference types inside a single generation.
Direct It With References
Pin a face, a look, a motion path, a camera move, or a voice using source material — the comfyui minimax h3 R2V node accepts up to 9 images, 3 videos, and 3 audio clips.
Crisp Text and Branding
Logos and spelled-out copy come out legible with the comfyui minimax h3 model, and plain-language instructions can describe how each reference relates to the rest.
Sage Attention Boost
Drop a Patch Sage Attention KJ node into the comfyui minimax h3 graph to roughly double throughput while barely touching output quality.
Resolution and Duration Grid
The comfyui minimax h3 Resolution Selector derives width and height from ratio and megapixels, snapped to the model's 32-pixel grid and 17-frame blocks at 24fps.
comfyui minimax h3 — Common Questions Answered
Answers to the questions people ask most about running the MiniMax H3 model inside ComfyUI.
What exactly is the comfyui minimax h3 workflow?
It is ComfyUI's built-in integration of MiniMax H3, MiniMax's general-purpose omni-modal model released with open weights. One forward pass turns text, pictures, footage, and audio references into video that already carries stereo sound.
How good is the output quality?
Clips can reach 2K at 24fps and run about 15 seconds. The native canvas uses a 768px short edge, tops out at 768x1344 pixels, and snaps dimensions to multiples of 32.
Which generation modes ship with it?
Three examples come in the comfyui minimax h3 template library: text-to-video (T2V), image-to-video (I2V) with optional first and last frame control, and reference-to-video (R2V) that locks a character, style, motion, camera, or voice.
Can it produce audio on its own?
It can. The comfyui minimax h3 model renders native stereo audio — speech, effects, and music — together with the picture in one pass, all synced inside a single MP4.
How do I get up and running?
Update ComfyUI to 0.30.0 or newer, open Template Library > Video, pick a comfyui minimax h3 workflow, then accept the pop-up that downloads models from the Hugging Face Comfy-Org/MiniMax-H3 repository.
Is there a way to make generation faster?
Install SageAttention plus the KJNodes custom nodes, then insert a Patch Sage Attention KJ node between the UNETLoader and BasicGuider in the comfyui minimax h3 graph to roughly double the speed.
Start Rendering with the comfyui minimax h3 Workflow
Bring MiniMax H3 to your own GPU inside ComfyUI: open weights, synced stereo sound, and full parameter access, with text-to-video, image-to-video, and reference-to-video graphs ready to load.
