MiniMax H3

MiniMax

MiniMax H3 is a general-purpose, omni-modal generative system. It supports unified understanding of multimodal contexts composed of text, images, video, and audio, and can generate video with native stereo audio at resolutions up to 2K and durations of up to 15 seconds.

Specs

Model

Family
MiniMax
Variant
H3
Parameters
33B
Weights
Open
License
MiniMax H3 Community License Agreement

Modalities

Input
Text Image Video Audio
Output
Video Audio

Dates

Announced
31 Jul 2026
Released
3 Aug 2026
MiniMax

MiniMax H3: An Open Model Breaking the Boundaries Between Tasks and Modalities

Today, we're officially launching MiniMax H3, a general-purpose omni-modal generation model. H3 can jointly understand multimodal contexts spanning text, images, video, and audio. It generates video with native stereo audio at up to 2K resolution and 15 seconds in length.

MiniMax
minimax.io

Keep up with the tools

An occasional email when notable AI dev tools and models land in the directory. No spam, unsubscribe anytime.