What MiniMax H3 Is Actually Good At
A hands-on local video study of source preservation, motion, native audio, Spectrum acceleration, reference control, and the limits of a 24 GB GPU.
In this article

The short answer
The curtain gave the models away. MiniMax H3 barely disturbed the finished advertisement around it. WAN 2.2 made the curtain billow, but also reconstructed more of the scene. LTX 2.3 moved fastest and changed the composition most. None was simply better. Each protected a different thing.
Across six finished layouts, H3 was our preferred result when preserving the source mattered most. It also generated stereo audio with the video. The tradeoff was clear: on an RTX 3090, H3 was the slowest route, and its restraint could look almost static when the brief called for bold movement.
The useful distinction
1. How we tested it
We ran a series of fixed-seed comparisons on an RTX 3090 with 24 GB of VRAM. Each comparison kept its source, prompt, requested duration, output shape, and seed fixed within that test. We saved the ComfyUI workflow, output, media metadata, and timing so that a surprising result could be checked instead of remembered.
- Six finished vertical advertisements animated with H3, WAN 2.2, and LTX 2.3.
- A 10, 12, 16, and 20-step H3 sweep on the same source and seed.
- Four fixed scenes comparing native 12, native 20, and conservative Spectrum 20.
- A generated-speech control comparing audio activity with visible mouth motion.
- A resolution and duration sweep that deliberately reached out-of-memory boundaries.
- Three reference-conditioning routes using the same two characters.
- A three-scene long-form chain comparing native quality with a six-step Turbo LoRA.
Automated metrics helped locate drift and timing changes, but visual and auditory review made the final decisions. Pixel motion can increase because a model is doing something useful, or because it is destroying the design. A single number cannot tell those cases apart.
This is also a comparison of products and workflows, not model weights in isolation. WAN used interpolation, LTX used its Director route, and H3 used its own audiovisual graph. A shared seed improves repeatability, but different models do not turn that seed into the same initial latent or sampling path.
2. Finished advertisements exposed H3's clearest advantage
Six finished vertical campaign layouts were resized to 704 by 1248. Each model received the same five-second target and a scene-specific prompt that locked typography, logo, crop, and graphic panels while requesting motion only in the photographed region. No overlay was reapplied afterward. This deliberately made text preservation part of the model test.
In our review, H3 preserved the layout best in all six cases. It also produced the least sampled frame change, which helps explain both its stability and its restrained feeling. WAN moved more. LTX moved most and finished much sooner, but rebuilt more of the composition.
| Workflow | Mean wall time | Layout-region drift | Frame-change proxy |
|---|---|---|---|
| MiniMax H3 | 14m 30s | 0.0018 | 0.0238 |
| WAN 2.2 + RIFE | 6m 41s | 0.0234 | 0.0487 |
| LTX 2.3 Director | 1m 39s | 0.0293 | 0.0842 |
The drift proxy is mean squared RGB change inside the left 62 percent of each source, where these layouts put most of their fixed graphics. It is not OCR and it does not measure text correctness. The frame-change proxy is mean absolute pixel change across five sampled moments. Lower values mean fewer changed pixels, which can indicate welcome stability or unwanted stillness. We used both to direct our eyes, not replace them.
Typography still needs a deterministic safety net
3. More steps helped, and Spectrum offered a qualified shortcut
Native 20 steps were our reference trajectory. Native 12 was faster but visibly weaker in the harder people and couple scenes. Conservative Spectrum kept a 20-step schedule while forecasting six of the expensive transformer evaluations. Across four fixed cases it averaged 647.2 seconds, about 91 seconds slower than native 12 and 248 seconds faster than native 20.

Decoded-video SSIM placed this Spectrum configuration closer to native 20 than native 12 in all four cases. SSIM measures structural pixel similarity, not narrative correctness, so we also watched and listened to every result. A more aggressive preset moved further from the native trajectory and reduced both motion and audio level.
Tested latency-quality compromise
schedule: 20 steps
native evaluations: 14
forecast evaluations: 6
sampler: RES multistep
reference: native 20Spectrum is an approximation
4. Native audio changed the value of the model
H3 predicts separate visual and audio latents inside one joint generation process. That lets visible action, ambience, effects, music, and speech influence the same sequence. It does not guarantee accurate words or perfect lip sync.
In this five-second control, a rough mouth-motion and audio-energy correlation produced the same 0.125-second offset for native 20 and the tested Spectrum run, while wall time fell from about nine minutes to six and a half. That number is a screening heuristic, not a lip-sync score. It compares frame differences in a mouth crop with audio RMS energy and can be fooled by a head turn, a blink, or background sound.
Listening still has to catch wrong words, stutter, extra voices, or an unnatural onset. A later continuation test produced exactly that kind of onset artifact even though the media file and overall waveform looked valid.
5. A 24 GB GPU has a real space-time boundary
Resolution and duration compete for the same memory. In this software build, the strongest reusable native spatial preset was 1536 by 864 for eight seconds. The strongest reusable duration preset was 1344 by 768 for ten seconds. Both retained native stereo audio. A 1600 by 896 eight-second preflight technically passed with only 70 MiB free, which is too fragile for a shared worker.
| Native request | Result | Peak VRAM | Quality wall time |
|---|---|---|---|
| 1536x864, 8s | Reusable pass | 23,198 MiB | 35m 27s |
| 1344x768, 10s | Reusable pass | 22,814 MiB | 34m 36s |
| 1344x768, 12s | Sampler OOM | 23,614 MiB at failure | Stopped in preflight |
| 1888x1056, 15s | Sampler OOM | Not a viable native shape | Stopped in preflight |
SageAttention moved the boundary in this setup, but it did not make 2 MP at 15 seconds fit. The remaining failures occurred in INT8 linear operations inside the transformer. Upscaling a smaller 15-second render can create a 2 MP delivery file, but that is not the same as native 2 MP detail.
This does not contradict MiniMax's broader claim of up to 2K output. The complete H3 system includes a Regenerate-2K stage, while the open H3 Base checkpoints generate at 768p and that official 2K regeneration module is not currently released for local use. Our external upscale produced a larger delivery file, not the same result as native sampling or the unreleased official regeneration stage.

6. Reference control is promising, but not settled
In the first two-character test, we passed the same references through three routes at the same seed, 960 by 544 resolution, eight seconds, and native 20-step sampling. The official Ref2VA route was soft and relatively static in this example. An experimental FL2VA route using the same reference node looked sharper.
The late-block hybrid kept a little more of the mage's hair curl and the gold material on his robe. It also cost 9.3 percent more wall time and about 506 MiB more peak VRAM than plain FL2VA. That was one seed, not enough to crown a new default.
Two follow-up subjects made the conclusion less tidy. On a character with a product, the hybrid had a small material advantage but no decisive identity gain. On a branded bicycle, both experimental routes approximated the lettering and neither was safe for client-facing brand detail. The official Ref2VA checkpoint remains the sensible general reference route until a broader comparison proves otherwise.
The practical lesson
7. Long-form video needs context management, not one giant render
In one three-scene test, native 20-step H3 produced the better reference-heavy chain but took 50.8 minutes at 960 by 544. The six-step LightX Turbo route was 2.64 times faster and failed semantically: one character changed into another, duplicates appeared, and an audio onset jumped at a join. Upscaling could not repair those failures.
A third-party context-loop workflow then carried 22 video frames and 22 audio frames between scenes and removed that overlap from both streams. It reduced the visual seam and eliminated the malformed vocal spike seen in a naive previous-video reference chain. It still could not force new dialogue to start at an exact timestamp.
One successful chain does not establish a universal recipe, but it supports a practical architecture: generate bounded shots, carry compact context, keep persistent identity references, checkpoint every accepted segment, and master long-form audio separately when exact chronology matters.
8. How to choose a route
H3 adds a useful high-fidelity option, but the model is too slow and too restrained to be the automatic answer to every brief. Start with the cost of failure. If a changed logo or face ruins the shot, test H3. If a still result ruins it, compare WAN. If iteration speed matters most, LTX may be the better creative tool.
- Keep exact typography deterministic when the design permits it.
- Test H3 first for completed designs and reference-critical shots.
- Compare WAN when stronger movement matters and LTX when speed matters.
- Treat fast Turbo reference chains as previews until identity and chronology pass review.
- Preflight risky space-time shapes before committing to a full 20-step render.
- Judge audio onset, dialogue accuracy, identity, and geometry separately.
The cleanest interface is not a wall of model switches. Ask whether the user wants to protect the source, maximize movement, or iterate quickly, then map that intent to a tested route. The model name is an implementation detail; the creative tradeoff is the useful choice.
Frequently asked questions
Is MiniMax H3 the best local video model for every job?
No. In this workflow-level comparison, H3 preserved finished layouts best. WAN 2.2 created stronger motion, while LTX 2.3 was much faster. The right choice depends on whether the shot values source fidelity, motion, or turnaround most.
Should finished advertisements still use deterministic text overlays?
Yes. H3 preserved existing text unusually well, but generated pixels are never an absolute typography guarantee. The safest production route still animates a clean visual base and reapplies exact logo and typography after motion.
What is the recommended H3 sampler setup from these tests?
Native 20 steps remained the reference in our tests. One conservative Spectrum configuration reduced runtime while staying closer to native 20 than native 12 by SSIM, but Spectrum is approximate and its implementation has continued to evolve. Validate the installed version with same-seed comparisons before using it for final work.
Can a 24 GB RTX 3090 generate 2 MP for 15 seconds natively?
Not with our tested open local workflow. A 15-second result required a smaller native render followed by an external upscale. That is different from native 2 MP detail and from MiniMax's complete hosted H3 system, which includes modules that are not all available as local checkpoints.
Sources
Primary model, implementation, and acceleration sources used alongside Movey's fixed-seed experiment records.
- Official MiniMax H3 model card documents the open FL2VA and Ref2VA checkpoints, joint audio-video architecture, supported durations, and the distinction between H3 Base and the complete 2K system.
- Official ComfyUI MiniMax H3 guide documents the local model modes, native stereo audio workflow, resolution guidance, and required ComfyUI version.
- Comfy-Org MiniMax H3 model repository provides the packaged local model assets used by the ComfyUI workflows.
- ComfyUI core documents native audiovisual generation support, model loading, offloading, quantization, and workflow execution.
- Spectrum for MiniMax H3 implements the forecasted transformer evaluations tested here and documents the method's approximation and audio-related limitations.
- Spectrum research implementation describes adaptive spectral feature forecasting for diffusion sampling acceleration.
- MiniMax H3 Contex Loop provides the bounded video and audio context mechanism used for the continuation follow-up.
Keep reading
Related articles

LTX 2.5 vs LTX 2.3: What Actually Improved
A visual LTX 2.5 and LTX 2.3 comparison covering prompt adherence, timed scenes, runtime compatibility, native audio, and practical upgrade limits.

Why AI Video Is So Much Heavier Than AI Images
Understand why AI video needs more GPU memory and compute than image generation, with tensor math, frame-resolution scaling, temporal attention, and practical cost controls.