Video generation
MiniMax H3 undercuts video rivals and plans open weights
MiniMax packs image, video and audio generation into one model, prices output at under a third of mainstream per-second rates, and plans to open the weights within days. How well the outputs hold up outside its own demos is the open question.
Emmanuel Fabrice Omgbwa Yasse AI-assisted
2026-07-30 · Last updated: 2026-08-02 · 4 min read

MiniMax is launching H3 as a bet on the economics of video generation. The model takes a mixed context of text, images, video and sound, and answers with footage that carries native stereo audio, up to 15 seconds at 2K. Per-second output is priced at under a third of mainstream models, and the weights are promised within days.
Rivals notice the second promise first. Video generation, MiniMax argues, still runs on closed models, and iteration speed and ecosystem openness have lagged the large language model world. Opening H3 applies the playbook that spread open LLMs far and cheap to a category that has so far resisted it, the same playbook behind China's open-weight gambit.
One model, no specialist stack
The design principle is unification. Earlier MiniMax systems, Hailuo 01 and Hailuo 02, were video models built up generation by generation. H3 collapses the specialist divisions, separate experts for text-to-image, editing, voice, effects, music and video-reference tasks, into one model trained on them together. MiniMax calls unification and generalization its first principle, and the payoff shows in the demo prompt: "Use the Hitchcock camera movement from video 1, make the person in image 2 sing, and take the voice from audio 3." One instruction, four input modalities, no routing between sub-models.
Early invite-testing feedback, per the company, shows commercial-grade results on instruction following, rendering text and brand elements, and video-to-video motion transfer, across advertising, e-commerce, product design, UI/UX and games. In pretraining, H3 mixes text-to-image, text-to-video with joint audio generation, multi-shot modeling, text-to-audio that treats voice, effects and music as one problem, and generalized reference and editing across image, audio and video pairs.
Where the low price comes from
The sharpest number in the release is the price. At 2K, output costs under a third of the per-second rate of mainstream models; at 768P, about half of mainstream 720P rates. No independent testing sits behind those figures yet:
| Claim | Detail (per MiniMax) |
|---|---|
| Default output | Native 2K, clips up to 15 seconds |
| Per-second price at 2K | Under 1/3 of mainstream models |
| Per-second price at 768P | About 1/2 of mainstream 720P |
| H3-VAE compression | 4x sequence length gain |
| Training throughput | Nearly 30% higher end to end |
| Caption pipeline | ~100K inference tokens per sample, ~4K caption tokens on average |
Two design choices make that price believable on paper. The H3-VAE tokenizer is a full rework of the previous generation's, and its higher compression buys a 4x sequence length gain that cuts training and inference cost. That compression is also why 2K is the default resolution. Upscaling has no dedicated super-resolution module; the base model regenerates its own low-resolution output in context, reusing the multimodal context to recover small text and details a conventional upscaler would have to guess. MiniMax is not the only lab attacking video-generation cost head on (Nvidia's hybrid-attention breakthrough).
Context understanding was a cost center of its own: most training samples consumed around 100K tokens of inference and produced, on average, about 4K tokens of caption. MiniMax calls the pipeline Contextual Omni Representation. It is where the model learns to follow instructions that mix modalities and to describe relationships inside a multimodal context.
The architecture notes admit a u-turn. MiniMax dropped the Hailuo-02 design that once gave it a significant edge because that design added complexity once task generalization became the goal. Multimodal context tripled sequence length variance and split understanding and generation workloads apart, so the team moved to a heterogeneous training setup with per-workload hardware tuning and sample-level load balancing. MiniMax claims nearly 30% higher end-to-end training throughput from it.
The open-weights bet
The openness date is vague, "in the coming days," and conditional on applicable law. MiniMax's stated reasons: push the open-source community forward, ease adaptation to domestic Chinese chips (compatibility with several existing domestic accelerators was considered from the design stage), and let users customize their own version. The open-weights fight had already drawn a letter from 41 rivals (41 rivals signed the same open-weight letter).
None of that is neutral. The nod to domestic Chinese chips, without naming any, signals where MiniMax expects its user base to grow.
What H3 admits it can't do yet
MiniMax lists its own limits. Model scale limits how completely some capabilities land, and fine detail in some scenes still needs work, with higher resolution and finer quality on the roadmap. A fuller tech report is promised, and the next H-series version is planned to fuse with the M-series models. In MiniMax's view, language is a generalizable, scalable computing system, and multimodality should stay tightly coupled to it.
Independent verification will take time; the release carries no benchmark figures. Benchmarks have a patchy record as a stand-in for real-world performance (three cases where they got it wrong). What is verifiable today is the strategy: a capable-sounding omni-modal model, priced as a discount, with weights promised to the open-source community. For a market built on closed models, that is a stress test. If H3 lands near its pitch, rival pricing will have to move. If it does not, open weights still lower the floor for everyone else.
- Source : MiniMax H3 undercuts video rivals and plans open weights — 2026-07-31
Get the tech essentials in 3 minutes every morning
One email, every weekday, with what actually matters in AI and tech.