MiniMax today released its third-generation video model, "H3," as open-weight. While it follows the previous Hailuo 01 and 02 models, this marks the first time the company has released a model as open-weight. It also features native support for ComfyUI from day one.

H3 is capable of omni-modal processing, accepting text, images, videos, or audio as input. The output consists of videos up to 2K resolution and 15 seconds in length. Audio is not handled in a separate process but is output as spatial audio within the same pass as the video generation.

A key feature is the integration of five separate tasks into a single model. Since actual workflows involve multiple modalities, the model itself handles cross-modal resolution if the relationships between inputs are described in the prompt.

Motion transfer is particularly important for graph-based workflows. Camera work and rhythm can be supplied from a reference video, while the subject and style can be obtained from a different source. This enables iterative refinement of shots.

The model is also designed for local LLM execution, optimized to run on RTX 3060 graphics cards.


Source: MiniMax H3 Day-0 Support in ComfyUI: Open Weights, Native Audio, and 2K Video(HN 334pt・94コメント) (HN Search (backfill), 2026-08-03)