Skip to content
Articles
TechnicalVideo ModelsComfyUIOptimizationProduction

What MiniMax H3 Is Actually Good At

A hands-on local video study of source preservation, motion, native audio, Spectrum acceleration, reference control, and long-form continuation.

Published Aug 12, 2026Updated Aug 17, 202616 min readDifficulty 4/5
In this article
Three final frames comparing MiniMax H3, WAN 2.2, and LTX 2.3 on a finished real-estate advertisement
One finished real-estate advertisement after five seconds of motion. From left to right: MiniMax H3, WAN 2.2, and LTX 2.3. H3 retained the layout most closely, WAN moved the curtain more strongly, and LTX rebuilt more of the composition.

The short answer

The curtain gave the models away. MiniMax H3 barely disturbed the finished advertisement around it. WAN 2.2 made the curtain billow, but also reconstructed more of the scene. LTX 2.3 moved fastest and changed the composition most. None was simply better. Each protected a different thing.

Across six finished layouts, H3 was our preferred result when preserving the source mattered most. It also generated stereo audio with the video. The tradeoff was clear: H3 was the slowest route, and its restraint could look almost static when the brief called for bold movement.

The useful distinction

Choose H3 when the input already contains pixels you cannot casually replace. Choose WAN when movement carries the shot. Choose LTX when turnaround and directed scene structure matter more than preserving every part of a finished layout. This conclusion applies to the tested workflows, not every possible configuration.

Production availability

MiniMax H3 is now available in Movey under written commercial authorization from MiniMax. See the MiniMax H3 model page for current workflows, production limits, commercial-use conditions, and the distinction between upstream H3 capabilities and Movey's current route.

1. How we tested it

We ran a series of fixed-seed comparisons on an RTX 3090 with 24 GB of VRAM. Each comparison kept its source, prompt, requested duration, output shape, and seed fixed within that test. We saved the ComfyUI workflow, output, media metadata, and timing so that a surprising result could be checked instead of remembered.

  • Six finished vertical advertisements animated with H3, WAN 2.2, and LTX 2.3.
  • A 10, 12, 16, and 20-step H3 sweep on the same source and seed.
  • Four fixed scenes comparing native 12, native 20, and conservative Spectrum 20.
  • A generated-speech control comparing audio activity with visible mouth motion.
  • Three reference-conditioning routes using the same two characters.
  • A three-scene long-form chain comparing native quality with a six-step Turbo LoRA.

Automated metrics helped locate drift and timing changes, but visual and auditory review made the final decisions. Pixel motion can increase because a model is doing something useful, or because it is destroying the design. A single number cannot tell those cases apart.

This is also a comparison of products and workflows, not model weights in isolation. WAN 2.2 used interpolation, the LTX result used our LTX 2.3 generation workflow, and MiniMax H3 used its matching generation graph. A shared seed improves repeatability, but different models do not turn that seed into the same initial latent or sampling path.

2. Finished advertisements exposed H3's clearest advantage

Six finished vertical campaign layouts were resized to the same portrait dimensions. Each model received the same five-second target and a scene-specific prompt that locked typography, logo, crop, and graphic panels while requesting motion only in the photographed region. No overlay was reapplied afterward. This deliberately made text preservation part of the model test.

MiniMax H3. The keys, glasses, architecture, logo, and copy remain close to the input while light, curtain, and coastal details move.
WAN 2.2. Motion is stronger, but the curtain, tabletop, and coastline are reconstructed more aggressively.
LTX 2.3. This was fastest, but the composition, text panel, keys, and crop drift more visibly.

In our review, H3 preserved the layout best in all six cases. It also changed the fewest pixels, which helps explain both its stability and its restrained feeling. WAN moved more. LTX moved most and finished much sooner, but rebuilt more of the composition.

WorkflowMean timeSource matchMotion
MiniMax H314m 30sClosestRestrained
WAN 2.2 + RIFE6m 41sMiddleActive
LTX 2.3 workflow1m 39sFurthestMost active

The times are averages across the six layouts. Source match and motion combine manual review with pixel-change checks, but the labels describe only this test set. Restrained motion can protect a layout or leave a shot feeling lifeless, which is why the videos matter more than a single score.

Typography still needs a deterministic safety net

H3 preserved generated pixels very well, but an advertisement cannot accept almost-correct legal copy, a changed price, or one damaged logo frame. Movey's safest ad route still animates a clean base and reapplies exact typography afterward. H3 makes baked-in designs viable when that split is not possible.

3. More steps helped, and Spectrum offered a qualified shortcut

Native 20 steps were our reference trajectory. Native 12 was faster but visibly weaker in the harder people and couple scenes. Conservative Spectrum kept a 20-step schedule while forecasting six of the expensive transformer evaluations. Across four fixed cases it averaged about eleven minutes, roughly four minutes faster than native 20 but one and a half minutes slower than native 12.

Play or scrub the shared five-second timeline. Native 20 is the reference. Watch the center woman's face, the hands around the flower bed, and the foreground haze: Native 12 departs more, while the conservative Spectrum run follows the reference more closely. Audio is omitted here so the visual comparison stays clear.

Decoded-video SSIM placed this Spectrum configuration closer to native 20 than native 12 in all four cases. SSIM measures structural pixel similarity, not narrative correctness, so we also watched and listened to every result. A more aggressive preset moved further from the native trajectory and reduced both motion and audio level.

Tested latency-quality compromise
schedule: 20 steps
native evaluations: 14
forecast evaluations: 6
sampler: RES multistep
reference: native 20

Spectrum is an approximation

These numbers belong to the extension version and preset we tested. The implementation has since changed its audio handling and defaults, and its own documentation warns that forecasting is not lossless. Native 20 remains the fidelity reference. Re-run same-seed comparisons before adopting Spectrum in a final pipeline.

4. Native audio changed the value of the model

H3 predicts separate visual and audio latents inside one joint generation process. That lets visible action, ambience, effects, music, and speech influence the same sequence. It does not guarantee accurate words or perfect lip sync.

Native 20-step control with generated voice, room tone, and visible mouth movement.
Conservative Spectrum control using the same image, prompt, seed, duration, and sentence.

The two clips had similar visible speech timing, while the tested Spectrum run finished about two and a half minutes sooner. That is useful, but it is not proof of accurate lip sync. Watch whether mouth movement follows the cadence, then listen for the actual words and any abrupt changes in the voice.

Listening still has to catch wrong words, stutter, extra voices, or an unnatural onset. A later continuation test produced exactly that kind of onset artifact even though the media file and overall waveform looked valid.

5. Reference control is promising, but not settled

In the first two-character test, we passed the same references through three routes with the same fixed settings and native 20-step sampling. The official Ref2VA route was soft and relatively static in this example. An experimental FL2VA route using the same reference node looked sharper.

Left: dedicated Ref2VA. Middle: FL2VA with references. Right: an experimental late-block FL2VA hybrid. Pause on the mage's hair curl and the gold material on his red robe.

The late-block hybrid kept a little more of the mage's hair curl and the gold material on his robe, but it was slower and the visual gain was small. That was one seed, not enough to crown a new default.

Two follow-up subjects made the conclusion less tidy. On a character with a product, the hybrid had a small material advantage but no decisive identity gain. On a branded bicycle, both experimental routes approximated the lettering and neither was safe for client-facing brand detail. The official Ref2VA checkpoint remains the sensible general reference route until a broader comparison proves otherwise.

The practical lesson

Sharpness and identity are different properties. A crisp frame can still contain the wrong face, material, or logo. Compare references at synchronized frames and judge identity, material, geometry, and text separately.

6. Long-form video needs context management, not one giant render

In one three-scene test, native 20-step H3 produced the better reference-heavy chain but took about 51 minutes. The six-step LightX Turbo route was roughly 2.6 times faster and failed semantically: one character changed into another, duplicates appeared, and an audio onset jumped at a join. Upscaling could not repair those failures.

Three nominal eight-second scenes generated with persistent references and bounded audiovisual context. The final assembly removes repeated context frames and preserves one Queen and one mage.

A third-party context-loop workflow then carried a short overlap of video and audio into each new scene and removed that overlap from the final assembly. It reduced the visual seam and eliminated the malformed vocal spike seen in a naive continuation. It still could not force new dialogue to start at an exact timestamp.

One successful chain does not establish a universal recipe, but it supports a practical architecture: generate bounded shots, carry compact context, keep persistent identity references, checkpoint every accepted segment, and master long-form audio separately when exact chronology matters.

7. How to choose a route

H3 adds a useful high-fidelity option, but the model is too slow and too restrained to be the automatic answer to every brief. Start with the cost of failure. If a changed logo or face ruins the shot, test H3. If a still result ruins it, compare WAN. If iteration speed matters most, LTX may be the better creative tool.

  • Keep exact typography deterministic when the design permits it.
  • Test H3 first for completed designs and reference-critical shots.
  • Compare WAN when stronger movement matters and LTX when speed matters.
  • Treat fast Turbo reference chains as previews until identity and chronology pass review.
  • Judge audio onset, dialogue accuracy, identity, and geometry separately.

The cleanest interface is not a wall of model switches. Ask whether the user wants to protect the source, maximize movement, or iterate quickly, then map that intent to a tested route. The model name is an implementation detail; the creative tradeoff is the useful choice.

Compare the two strongest routes directly in MiniMax H3 vs WAN 2.2, then test a source frame in Movey's image-to-video workflow.

Frequently asked questions

Can I use MiniMax H3 commercially through Movey?

Yes. Movey has written commercial authorization from MiniMax to provide H3 and H3 Works. Use remains subject to your plan, applicable law, third-party rights, Movey's terms, MiniMax's model terms, and supported-territory restrictions.

Is MiniMax H3 the best local video model for every job?

No. In this workflow-level comparison, H3 preserved finished layouts best. WAN 2.2 created stronger motion, while LTX 2.3 was much faster. The right choice depends on whether the shot values source fidelity, motion, or turnaround most.

Should finished advertisements still use deterministic text overlays?

Yes. H3 preserved existing text unusually well, but generated pixels are never an absolute typography guarantee. The safest production route still animates a clean visual base and reapplies exact logo and typography after motion.

What is the recommended H3 sampler setup from these tests?

Native 20 steps remained the reference in our tests. One conservative Spectrum configuration reduced runtime while staying closer to native 20 than native 12 by SSIM, but Spectrum is approximate and its implementation has continued to evolve. Validate the installed version with same-seed comparisons before using it for final work.

Does MiniMax H3 generate audio with the video?

Yes. H3 jointly generates visual and audio latents, so dialogue, ambience, sound effects, and music can emerge with the picture. The result still needs listening review because generated words, timing, and audio onset can be wrong.

Sources

Primary model, implementation, and acceleration sources used to interpret the fixed-seed comparisons.

Keep reading

Related articles

All guides
Fixed-seed comparison frames from LTX 2.3 and the full LTX 2.5 model stack
TechnicalVideo Models

LTX 2.5 vs 2.3: Quality, Speed & Audio Test

LTX 2.5 vs 2.3 tested with fixed seeds: compare prompt adherence, scene continuity, speed, native audio, and the full ComfyUI upgrade path.

15 min readDifficulty 4/5
What MiniMax H3 Is Actually Good At | Movey