MiniMax releases MiniMax H3

03-08-2026

MiniMax H3 is MiniMax's omni-modal generation model, released on July 31, 2026, generating up to 15 seconds of 2K video with native stereo audio from a unified text, image, video and audio context, with open weights announced.

Written by:

Jorick van Weelie

Marketing Lead at DataNorth | AI Enthusiast & Tech Storyteller

minimax releases minimax h3
Sign up for our Newsletter

Published: August 03, 2026

MiniMax released MiniMax H3 on July 31, 2026, a general-purpose omni-modal generation model that reads text, images, video and audio as a single context and returns video with native stereo sound at up to 2K resolution and 15 seconds in length. MiniMax H3 is live in the MiniMax platform API under the model ID MiniMax-H3 and in the consumer Hailuo AI app, and MiniMax has said it plans to publish the model weights. At 2K, MiniMax prices H3 at less than a third of the per-second cost of mainstream video models.

What did MiniMax release with MiniMax H3?

MiniMax H3 is the third generation of the company’s video model line, following Hailuo 01 and Hailuo 02. Where those models split generation into separate expert systems for text-to-image, text-to-video, image-to-video, first-and-last-frame, subject reference, motion reference and video editing, H3 folds all of them into a single pre-training paradigm. Reference and editing relationships are expressed in natural language rather than selected from a fixed task list, which is what MiniMax means when it calls H3 a general-purpose model rather than a video model.

Audio is generated jointly with the video rather than added afterwards. Voice, sound effects and music are modelled together instead of as separate domains, and all audio output is native stereo. The model also supports native multi-shot generation. MiniMax states that H3 handles a prompt such as referencing camera movement from one input video, a character from a second input image and a vocal track from a third input audio file, with the relationship between those inputs described in words.

MiniMax H3 technical specs and architecture

MiniMax H3 outputs video at up to 2K resolution, in clips of up to 15 seconds, with native stereo audio. Four components carry the release. H3-Contextual Omni Representation extends captioning to describe the relationship between the input context and the target video, using a dedicated full-modality understanding pipeline that takes roughly 100,000 tokens of inference per source and distils it to an average of about 4,000 tokens. H3-VAE is a rebuilt tokenizer whose higher compression ratio delivers a four times gain in effective sequence length and is the technology behind native 2K support.

H3-Omni Transformer is the architecture. MiniMax set aside the Hailuo 02 architecture for this release, arguing that its efficiency advantages added complexity that worked against task generalisation. Because multimodal context tripled the variance in sequence length, MiniMax adopted a training architecture that separates understanding and generation workloads and tunes hardware utilisation for each, which lifted end-to-end training throughput by close to 30 percent. H3-In-Context Regeneration replaces a conventional super-resolution module: for 2K output the base model regenerates its own low-resolution result in context, drawing on the original multimodal input again to recover small text and fine detail.

MiniMax H3 pricing and how it compares to closed video models

MiniMax offers 2K resolution by default and prices H3 at less than a third of the per-second cost of mainstream video models at that resolution. At 768p, MiniMax puts the price at less than half that of mainstream models at 720p. The company frames this price-performance position as the direct result of the H3-VAE compression work rather than a promotional rate.

The competitive argument MiniMax makes is about openness rather than raw quality. Video generation has been dominated by closed models with slower iteration and a less open ecosystem than large language models, and MiniMax says hardware compatibility was a design consideration for H3 from the earliest stages so that the model can run across a broader range of AI hardware once weights are published. MiniMax reports that early testing shows H3 performing well on instruction following, accurate text and brand rendering, and video-to-video motion transfer, and positions it for advertising, branding, e-commerce, product design, UI and UX, and gaming.

MiniMax H3 availability and open weights

MiniMax H3 is available now through the MiniMax platform API under the model ID MiniMax-H3 and in the Hailuo AI consumer app. MiniMax has stated that it plans to open the model weights in the coming days, subject to applicable laws and regulations. At the time of writing, no Hugging Face repository or model card had been published, so teams planning to self-host should treat the open-weights timeline as announced rather than delivered.

MiniMax has also said a full H3 technical report will follow. For the next generation of the H series, the company lists three priorities: integrating capabilities from its M-series language models to strengthen multimodal understanding, scaling the model size, and improving visual fidelity at higher resolutions.

Full details are in the official MiniMax H3 announcement.

Add DataNorth AI to your Google favorites