Skip to content
Articles
TechnicalVideo ModelsComfyUIOptimizationProduction

What MiniMax H3 Is Actually Good At

A hands-on local video study of source preservation, motion, native audio, Spectrum acceleration, reference control, and the limits of a 24 GB GPU.

Published Aug 12, 202619 min readDifficulty 4/5
In this article
Three final frames comparing MiniMax H3, WAN 2.2, and LTX 2.3 on a finished real-estate advertisement
One finished real-estate advertisement after five seconds of motion. From left to right: MiniMax H3, WAN 2.2, and LTX 2.3. H3 retained the layout most closely, WAN moved the curtain more strongly, and LTX rebuilt more of the composition.

The short answer

The curtain gave the models away. MiniMax H3 barely disturbed the finished advertisement around it. WAN 2.2 made the curtain billow, but also reconstructed more of the scene. LTX 2.3 moved fastest and changed the composition most. None was simply better. Each protected a different thing.

Across six finished layouts, H3 was our preferred result when preserving the source mattered most. It also generated stereo audio with the video. The tradeoff was clear: on an RTX 3090, H3 was the slowest route, and its restraint could look almost static when the brief called for bold movement.

The useful distinction

Choose H3 when the input already contains pixels you cannot casually replace. Choose WAN when movement carries the shot. Choose LTX when turnaround and directed scene structure matter more than preserving every part of a finished layout. This conclusion applies to the tested workflows, not every possible configuration.

1. How we tested it

We ran a series of fixed-seed comparisons on an RTX 3090 with 24 GB of VRAM. Each comparison kept its source, prompt, requested duration, output shape, and seed fixed within that test. We saved the ComfyUI workflow, output, media metadata, and timing so that a surprising result could be checked instead of remembered.

  • Six finished vertical advertisements animated with H3, WAN 2.2, and LTX 2.3.
  • A 10, 12, 16, and 20-step H3 sweep on the same source and seed.
  • Four fixed scenes comparing native 12, native 20, and conservative Spectrum 20.
  • A generated-speech control comparing audio activity with visible mouth motion.
  • A resolution and duration sweep that deliberately reached out-of-memory boundaries.
  • Three reference-conditioning routes using the same two characters.
  • A three-scene long-form chain comparing native quality with a six-step Turbo LoRA.

Automated metrics helped locate drift and timing changes, but visual and auditory review made the final decisions. Pixel motion can increase because a model is doing something useful, or because it is destroying the design. A single number cannot tell those cases apart.

This is also a comparison of products and workflows, not model weights in isolation. WAN used interpolation, LTX used its Director route, and H3 used its own audiovisual graph. A shared seed improves repeatability, but different models do not turn that seed into the same initial latent or sampling path.

2. Finished advertisements exposed H3's clearest advantage

Six finished vertical campaign layouts were resized to 704 by 1248. Each model received the same five-second target and a scene-specific prompt that locked typography, logo, crop, and graphic panels while requesting motion only in the photographed region. No overlay was reapplied afterward. This deliberately made text preservation part of the model test.

MiniMax H3. The keys, glasses, architecture, logo, and copy remain close to the input while light, curtain, and coastal details move.
WAN 2.2. Motion is stronger, but the curtain, tabletop, and coastline are reconstructed more aggressively.
LTX 2.3. This was fastest, but the composition, text panel, keys, and crop drift more visibly.

In our review, H3 preserved the layout best in all six cases. It also produced the least sampled frame change, which helps explain both its stability and its restrained feeling. WAN moved more. LTX moved most and finished much sooner, but rebuilt more of the composition.

WorkflowMean wall timeLayout-region driftFrame-change proxy
MiniMax H314m 30s0.00180.0238
WAN 2.2 + RIFE6m 41s0.02340.0487
LTX 2.3 Director1m 39s0.02930.0842

The drift proxy is mean squared RGB change inside the left 62 percent of each source, where these layouts put most of their fixed graphics. It is not OCR and it does not measure text correctness. The frame-change proxy is mean absolute pixel change across five sampled moments. Lower values mean fewer changed pixels, which can indicate welcome stability or unwanted stillness. We used both to direct our eyes, not replace them.

Typography still needs a deterministic safety net

H3 preserved generated pixels very well, but an advertisement cannot accept almost-correct legal copy, a changed price, or one damaged logo frame. Movey's safest ad route still animates a clean base and reapplies exact typography afterward. H3 makes baked-in designs viable when that split is not possible.

3. More steps helped, and Spectrum offered a qualified shortcut

Native 20 steps were our reference trajectory. Native 12 was faster but visibly weaker in the harder people and couple scenes. Conservative Spectrum kept a 20-step schedule while forecasting six of the expensive transformer evaluations. Across four fixed cases it averaged 647.2 seconds, about 91 seconds slower than native 12 and 248 seconds faster than native 20.

Sampled Garden Grove frames from native 12, native 20, and conservative Spectrum 20 H3 runs
Top to bottom: native 12, native 20, and the tested conservative Spectrum 20 configuration. Look at hands, faces, local movement, and gradual exposure rather than judging only the first frame.

Decoded-video SSIM placed this Spectrum configuration closer to native 20 than native 12 in all four cases. SSIM measures structural pixel similarity, not narrative correctness, so we also watched and listened to every result. A more aggressive preset moved further from the native trajectory and reduced both motion and audio level.

Tested latency-quality compromise
schedule: 20 steps
native evaluations: 14
forecast evaluations: 6
sampler: RES multistep
reference: native 20

Spectrum is an approximation

These numbers belong to the extension version and preset we tested. The implementation has since changed its audio handling and defaults, and its own documentation warns that forecasting is not lossless. Native 20 remains the fidelity reference. Re-run same-seed comparisons before adopting Spectrum in a final pipeline.

4. Native audio changed the value of the model

H3 predicts separate visual and audio latents inside one joint generation process. That lets visible action, ambience, effects, music, and speech influence the same sequence. It does not guarantee accurate words or perfect lip sync.

Native 20-step control with generated voice, room tone, and visible mouth movement.
Conservative Spectrum control using the same image, prompt, seed, duration, and sentence.

In this five-second control, a rough mouth-motion and audio-energy correlation produced the same 0.125-second offset for native 20 and the tested Spectrum run, while wall time fell from about nine minutes to six and a half. That number is a screening heuristic, not a lip-sync score. It compares frame differences in a mouth crop with audio RMS energy and can be fooled by a head turn, a blink, or background sound.

Listening still has to catch wrong words, stutter, extra voices, or an unnatural onset. A later continuation test produced exactly that kind of onset artifact even though the media file and overall waveform looked valid.

5. A 24 GB GPU has a real space-time boundary

Resolution and duration compete for the same memory. In this software build, the strongest reusable native spatial preset was 1536 by 864 for eight seconds. The strongest reusable duration preset was 1344 by 768 for ten seconds. Both retained native stereo audio. A 1600 by 896 eight-second preflight technically passed with only 70 MiB free, which is too fragile for a shared worker.

Native requestResultPeak VRAMQuality wall time
1536x864, 8sReusable pass23,198 MiB35m 27s
1344x768, 10sReusable pass22,814 MiB34m 36s
1344x768, 12sSampler OOM23,614 MiB at failureStopped in preflight
1888x1056, 15sSampler OOMNot a viable native shapeStopped in preflight

SageAttention moved the boundary in this setup, but it did not make 2 MP at 15 seconds fit. The remaining failures occurred in INT8 linear operations inside the transformer. Upscaling a smaller 15-second render can create a 2 MP delivery file, but that is not the same as native 2 MP detail.

This does not contradict MiniMax's broader claim of up to 2K output. The complete H3 system includes a Regenerate-2K stage, while the open H3 Base checkpoints generate at 768p and that official 2K regeneration module is not currently released for local use. Our external upscale produced a larger delivery file, not the same result as native sampling or the unreleased official regeneration stage.

Frame from the 15-second native MiniMax H3 capacity run
The 15-second native source used 960 by 544 before a separate delivery upscale. A large output container does not retroactively create native facial detail.

6. Reference control is promising, but not settled

In the first two-character test, we passed the same references through three routes at the same seed, 960 by 544 resolution, eight seconds, and native 20-step sampling. The official Ref2VA route was soft and relatively static in this example. An experimental FL2VA route using the same reference node looked sharper.

Left: dedicated Ref2VA. Middle: traditional FL2VA with references. Right: FL2VA with reference AdaLN blocks 30 to 49. Pause on the mage's hair curl and the gold material on his red robe.

The late-block hybrid kept a little more of the mage's hair curl and the gold material on his robe. It also cost 9.3 percent more wall time and about 506 MiB more peak VRAM than plain FL2VA. That was one seed, not enough to crown a new default.

Two follow-up subjects made the conclusion less tidy. On a character with a product, the hybrid had a small material advantage but no decisive identity gain. On a branded bicycle, both experimental routes approximated the lettering and neither was safe for client-facing brand detail. The official Ref2VA checkpoint remains the sensible general reference route until a broader comparison proves otherwise.

The practical lesson

Sharpness and identity are different properties. A crisp frame can still contain the wrong face, material, or logo. Compare references at synchronized frames and judge identity, material, geometry, and text separately.

7. Long-form video needs context management, not one giant render

In one three-scene test, native 20-step H3 produced the better reference-heavy chain but took 50.8 minutes at 960 by 544. The six-step LightX Turbo route was 2.64 times faster and failed semantically: one character changed into another, duplicates appeared, and an audio onset jumped at a join. Upscaling could not repair those failures.

Three nominal eight-second scenes generated with persistent references and bounded audiovisual context. The final assembly removes repeated context frames and preserves one Queen and one mage.

A third-party context-loop workflow then carried 22 video frames and 22 audio frames between scenes and removed that overlap from both streams. It reduced the visual seam and eliminated the malformed vocal spike seen in a naive previous-video reference chain. It still could not force new dialogue to start at an exact timestamp.

One successful chain does not establish a universal recipe, but it supports a practical architecture: generate bounded shots, carry compact context, keep persistent identity references, checkpoint every accepted segment, and master long-form audio separately when exact chronology matters.

8. How to choose a route

H3 adds a useful high-fidelity option, but the model is too slow and too restrained to be the automatic answer to every brief. Start with the cost of failure. If a changed logo or face ruins the shot, test H3. If a still result ruins it, compare WAN. If iteration speed matters most, LTX may be the better creative tool.

  • Keep exact typography deterministic when the design permits it.
  • Test H3 first for completed designs and reference-critical shots.
  • Compare WAN when stronger movement matters and LTX when speed matters.
  • Treat fast Turbo reference chains as previews until identity and chronology pass review.
  • Preflight risky space-time shapes before committing to a full 20-step render.
  • Judge audio onset, dialogue accuracy, identity, and geometry separately.

The cleanest interface is not a wall of model switches. Ask whether the user wants to protect the source, maximize movement, or iterate quickly, then map that intent to a tested route. The model name is an implementation detail; the creative tradeoff is the useful choice.

Frequently asked questions

Is MiniMax H3 the best local video model for every job?

No. In this workflow-level comparison, H3 preserved finished layouts best. WAN 2.2 created stronger motion, while LTX 2.3 was much faster. The right choice depends on whether the shot values source fidelity, motion, or turnaround most.

Should finished advertisements still use deterministic text overlays?

Yes. H3 preserved existing text unusually well, but generated pixels are never an absolute typography guarantee. The safest production route still animates a clean visual base and reapplies exact logo and typography after motion.

What is the recommended H3 sampler setup from these tests?

Native 20 steps remained the reference in our tests. One conservative Spectrum configuration reduced runtime while staying closer to native 20 than native 12 by SSIM, but Spectrum is approximate and its implementation has continued to evolve. Validate the installed version with same-seed comparisons before using it for final work.

Can a 24 GB RTX 3090 generate 2 MP for 15 seconds natively?

Not with our tested open local workflow. A 15-second result required a smaller native render followed by an external upscale. That is different from native 2 MP detail and from MiniMax's complete hosted H3 system, which includes modules that are not all available as local checkpoints.

Sources

Primary model, implementation, and acceleration sources used alongside Movey's fixed-seed experiment records.

  • Official MiniMax H3 model card documents the open FL2VA and Ref2VA checkpoints, joint audio-video architecture, supported durations, and the distinction between H3 Base and the complete 2K system.
  • Official ComfyUI MiniMax H3 guide documents the local model modes, native stereo audio workflow, resolution guidance, and required ComfyUI version.
  • Comfy-Org MiniMax H3 model repository provides the packaged local model assets used by the ComfyUI workflows.
  • ComfyUI core documents native audiovisual generation support, model loading, offloading, quantization, and workflow execution.
  • Spectrum for MiniMax H3 implements the forecasted transformer evaluations tested here and documents the method's approximation and audio-related limitations.
  • Spectrum research implementation describes adaptive spectral feature forecasting for diffusion sampling acceleration.
  • MiniMax H3 Contex Loop provides the bounded video and audio context mechanism used for the continuation follow-up.

Keep reading

Related articles

All guides
Fixed-seed comparison frames from LTX 2.3 and the full LTX 2.5 model stack
TechnicalVideo Models

LTX 2.5 vs LTX 2.3: What Actually Improved

A visual LTX 2.5 and LTX 2.3 comparison covering prompt adherence, timed scenes, runtime compatibility, native audio, and practical upgrade limits.

15 min readDifficulty 4/5
Diagram showing image frames stacking into an increasingly large space-time video payload
TechnicalVideo Models

Why AI Video Is So Much Heavier Than AI Images

Understand why AI video needs more GPU memory and compute than image generation, with tensor math, frame-resolution scaling, temporal attention, and practical cost controls.

14 min readDifficulty 3/5
What MiniMax H3 Is Actually Good At | Movey