Feedback
AI Ad Video Example
Loading...
comfyui minimax h3
Generate open-weight video with synchronized sound from text, images, or reference footage using the comfyui minimax h3 workflow — up to 2K at 24fps.
All Tools
Discover our comprehensive AI-powered animation toolkit
MiniMax H3
MiniMax H3 AI Video Generator
Seedance 2.5
The Future of AI Video Is Here.
FLUX 3 Video Generator

Kling 3.0
Next-Gen AI Video Generator
Grok Video Generator
Create Videos from Text or Images with AI
MiniMax H3 video generator
MiniMax H3 AI Video Generator

Free GPT Image 2.5
Next-generation image creator

Nano Banana2
Best Image Generator
The Benefits of Running the comfyui minimax h3 Workflow Locally
By installing the comfyui minimax h3 workflow, you gain access to MiniMax's open-weight, omni-modal generation model inside ComfyUI. The workflow reads text, images, video, and audio as one shared context, then produces video and synchronized stereo audio — dialogue, sound effects, and music — in a single forward pass. Clips reach roughly 15 seconds at up to 2K resolution and 24fps, while every node remains configurable.
- Audio Built Into Every ClipSpeech, sound effects, and musical score appear in the same MP4 as the picture, locked together by the comfyui minimax h3 workflow in a single rendering pass.
- Full Local CustomizationBecause the model is open weight, you can host the comfyui minimax h3 workflow on your own machine and adjust resolution, runtime, and all diffusion settings without hitting API quotas.
- Combine Text, Images, and MoreFeed the comfyui minimax h3 nodes a prompt plus image, video, or audio references to keep a character, aesthetic, action, camera path, or voice consistent across the generated clip.
Three Steps to Start the comfyui minimax h3 Workflow
Follow this quick walkthrough to produce your first open-weight video with sound through the comfyui minimax h3 workflow in just three steps.
Core Strengths of the comfyui minimax h3 Workflow
The comfyui minimax h3 setup bundles three official ComfyUI examples, open-weight multimodal generation, synchronized stereo audio, reference-based control, and an optional Sage Attention speedup — everything you need for local video production.
Three Ready-to-Run ComfyUI Templates
Inside the comfyui minimax h3 library you will find text-to-video, image-to-video, and reference-to-video sample graphs, with each template dedicated to a distinct creation mode.
A Single Context for Every Input Type
The comfyui minimax h3 model processes text, visuals, footage, and sound as one unified context, so you can mix all reference types into a single generation.
Keep Characters and Styles Consistent
Use reference files to hold onto a character, visual style, movement, camera motion, or voice — the comfyui minimax h3 R2V node accepts up to 9 images, 3 videos, and 3 audio clips.
Sharp Text Rendering and Brand Consistency
The comfyui minimax h3 model reproduces legible text and brand elements accurately, and follows natural-language instructions that define how references relate to each other.
Approximately Double Your Render Speed
Insert the Patch Sage Attention KJ node into the comfyui minimax h3 graph to nearly halve render time while keeping quality essentially unchanged.
Precise Resolution and Duration Controls
The comfyui minimax h3 Resolution Selector derives width and height from aspect ratio plus megapixels, then rounds to the model's 32-pixel multiples and 17-frame-per-block timing at 24fps.
Got Questions About the comfyui minimax h3 Workflow?
Straightforward answers for running the MiniMax H3 model under ComfyUI with the comfyui minimax h3 workflow.
What exactly does the comfyui minimax h3 workflow do?
This is the official ComfyUI integration for MiniMax H3, an open-weight, general-purpose omni-modal model. With this workflow, you can produce video plus native stereo audio from text, images, video, or audio inputs all at once.
How high can the output resolution go?
Video from the comfyui minimax h3 workflow reaches up to 2K at 24fps for around 15 seconds. Internally, the canvas uses a 768px short edge, limits to 768x1344, and snaps to 32-pixel multiples.
Can I create video from text, image, or references?
Yes — the comfyui minimax h3 templates cover text-to-video (T2V), image-to-video (I2V) with optional first/last frame control, and reference-to-video (R2V) for preserving character, style, motion, camera, or voice.
Will the workflow create sound too?
Absolutely. The comfyui minimax h3 model renders stereo audio — speech, sound effects, and music — together with the visuals, outputting everything synced inside one MP4 file.
What do I need to do to begin?
Ensure ComfyUI is version 0.30.0 or later, head to Template Library > Video, choose a comfyui minimax h3 workflow, and follow the on-screen prompts to download weights from the Hugging Face Comfy-Org/MiniMax-H3 repository.
Is there a way to make rendering noticeably faster?
You can — install SageAttention and KJNodes, then insert a Patch Sage Attention KJ node between the UNETLoader and BasicGuider in the comfyui minimax h3 graph. That roughly doubles generation throughput.
Begin Generating with the comfyui minimax h3 Workflow Today
Kick off local MiniMax H3 generation inside ComfyUI using the comfyui minimax h3 workflow — stereo audio, open weights, and total parameter control are built in for T2V, I2V, and R2V projects.
