2026/08/07· Last verified 2026/08/07

MiniMax H3 vs Wan 2.7 and Wan 2.2 Open Weights

Compare MiniMax H3 with hosted Wan 2.7 and Wan 2.2 open weights across multimodal control, audio, local VRAM, ComfyUI, quality, and production cost.

MiniMax H3 vs Wan 2.7 and Wan 2.2 Open Weights cover

Quick answer: Compare MiniMax H3 with Wan 2.7 when evaluating current hosted multimodal generation and editing. Compare H3 released resources with Wan 2.2 when the question is open weights, ComfyUI, VRAM, or local control. Treating hosted Wan 2.7 and open Wan 2.2 as the same product produces a misleading result.

Wan version status

Alibaba Cloud Model Studio documents Wan 2.7 as the current hosted family for text-to-video, image-to-video, reference-to-video, continuation, and instruction-based editing. Its image-to-video interface accepts text, image, audio, and video material and supports first-frame, first-and-last-frame, and continuation tasks. The official Wan open-weight GitHub organization, however, still exposes Wan 2.2 as the relevant downloadable generation repository.

That difference changes the test. Wan 2.7 belongs in a current hosted API comparison against the H3 online workflow. Wan 2.2 belongs in a reproducible local comparison with exact checkpoints, precision, node graph, and hardware. This page covers both because the unversioned search “MiniMax H3 vs Wan” can mean either intent.

A MiniMax H3 vs Wan comparison becomes confusing when “Wan” is treated as one file. The official Wan 2.2 repository includes several model variants with different sizes, inputs, resolutions, and hardware paths. H3 also differs between its official released resources, community quantizations, and the complete hosted workflow. Any useful comparison must name the checkpoint, precision, resolution, frame count, node graph, and machine.

Comparison at a glance

AreaMiniMax H3Wan 2.2 ecosystem
Core focusUnified multimodal video with synchronized audioFamily of open video models for several generation tasks
InputsText, images, video, and audio in documented H3 modesText/image-to-video and specialized variants such as speech-to-video
Hosted output768P or direct 2K, 4–15 secondsHosted options vary; official local variants document their own output settings
Local availabilityOfficial model repository plus community workflowsOfficial code and several model variants
AudioNative synchronized audio in H3 workflowAudio behavior depends on the selected Wan variant and external components
HardwareDepends on H3 checkpoint, precision, quantization, and offloadingDepends strongly on 5B/14B variant, task, precision, and offloading
Best comparisonExact H3 workflow versus exact Wan checkpointExact Wan checkpoint versus exact H3 workflow

Do not compare a small Wan text/image model to a larger specialized H3 workflow without stating the difference. Model size and task scope affect memory, speed, and capability.

Local installation and ComfyUI

Both ecosystems attract ComfyUI users. H3 local workflows may combine an H3 checkpoint, video VAE, text or vision-language encoder, scheduler, attention optimization, and community nodes. The H3 ComfyUI guide and VRAM guide explain why one memory number is unreliable.

Wan 2.2's official repository documents multiple pipelines and distributed options. The TI2V-5B variant supports text/image-to-video generation and is positioned as a more accessible model than 14B variants, while specialized speech-to-video setups add audio and pose-related components. A ComfyUI workflow may simplify the graph but does not erase checkpoint differences.

For a fair local benchmark, record GPU, VRAM, system RAM, operating system, driver, PyTorch, attention backend, checkpoint revision, precision, quantization, offload mode, resolution, frames, steps, sampler, warm-up, and total generation time.

Multimodal control

H3's distinctive workflow is the ability to state relationships between different reference modalities: appearance from an image, movement from video, and voice or ambience from audio. It also offers first/last-frame control as another generation mode.

Wan control depends on the exact model. A text/image-to-video checkpoint is not equivalent to H3 reference-to-video, while a speech-to-video checkpoint serves a more specialized performance task. Wan may be preferable when its dedicated model exactly matches the job; H3 may be preferable when several reference types must be coordinated in one directing brief.

Audio workflow

H3 generates synchronized stereo audio in the documented hosted result. That can reduce the need for separate sound-effect, ambience, or dialogue passes, although generated audio still needs review.

Wan 2.2 includes specialized speech-to-video research and workflows, but audio capability should be attributed to the exact variant. Some common text/image-to-video pipelines produce video without the same integrated soundtrack behavior. If a Wan pipeline uses CosyVoice or another component, disclose it rather than crediting every Wan checkpoint with the complete audio stack.

Test three cases: visible dialogue, a physical Foley action, and an ambience-heavy scene. Compare not only whether sound exists but whether it is synchronized, stable, editable, and appropriate.

VRAM and speed

Community discussions often reduce this comparison to “which runs on my GPU.” The answer changes with quantization and resolution. A smaller or aggressively quantized model may load in less memory but lose quality, slow down through offloading, or require unsupported nodes.

Measure peak VRAM, peak system RAM, cold-start time, warm generation time, and failures. Do not publish a single result as a universal minimum. If one workflow uses a pruned INT8 checkpoint and the other full precision, label that clearly.

Quality and accepted-shot cost

Local generation has no per-request provider fee, but it is not free. Include hardware depreciation, electricity, storage, setup, failed runs, and operator time. Hosted H3 has a visible credit cost and removes local maintenance. Heavy sustained workloads may favor local economics; occasional 2K jobs may favor hosted access.

For quality, test motion, identity, hands, text, camera, reference transfer, and sound. Score accepted outputs, not cherry-picked frames. The real-world test framework provides a reusable rubric.

Same-prompt benchmark

Eight-second documentary wildlife shot at blue hour. A red fox walks across a shallow frozen stream, pauses when the ice cracks softly, then looks toward camera as snow moves through the frame. Low tracking camera at shoulder height, realistic weight and paw contact, stable anatomy and fur markings, natural breath, restrained color, no cuts, no text. Include wind, paws on ice, one subtle crack, and no music when the workflow supports synchronized audio.

Use the same starting image for image-to-video variants. Match resolution and frame count as closely as possible. Generate four outputs with a fixed retry policy. Record local machine data and any separate audio stage.

Who should choose each?

Choose H3 when coordinated image/video/audio references, synchronized soundtrack, first/last frames, hosted direct 2K, or an H3-specific workspace matters. Choose Wan 2.2 when an official Wan variant fits the task, local control is required, or the team already maintains a validated Wan pipeline. Test both when local production economics and multimodal control are equally important.

FAQ

What is the current Wan version?

Wan 2.7 is the current hosted Alibaba Model Studio family, while Wan 2.2 remains the relevant official open-weight baseline.

Which needs less VRAM?

It depends on checkpoint size, precision, quantization, resolution, frames, and offloading. There is no responsible universal answer without a configuration.

Why are Wan 2.7 and Wan 2.2 both included?

They answer different questions: hosted current-model access versus open local experimentation. Results should never be merged under one Wan label.

Which is cheaper?

Hosted H3 uses credits; local Wan and H3 use owned hardware and operator time. Compare total accepted-shot cost for the expected workload.

Primary sources

Last verified August 7, 2026. This independent comparison is not affiliated with MiniMax or Alibaba.

Newsletter

Join the community

Subscribe to our newsletter for the latest news and updates