
Wan 3.0 vs MiniMax H3: Video Inputs, Limits and Cost
Compare Wan 3.0 vs MiniMax H3 by clip length, reference inputs, native audio and reAPI prices. See 2K cost estimates and input charges before choosing a model.
Choose Wan 3.0 for a continuous shot longer than fifteen seconds, document or webpage input, or a request that needs its larger per-type reference limits. Choose MiniMax H3 when a four-to-fifteen-second shot fits and its lower listed 2K output rate matters. Both generate video with audio; the decision is about the job and its constraints.[1][2][3][4]
This Wan 3.0 vs MiniMax H3 comparison uses public reAPI documentation and rendered price tables checked on September 14, 2026, with MiniMax's own release materials for the local-model distinction. It does not claim a measured image-quality or speed winner. Start with the Wan 3.0 or MiniMax H3 page after choosing the input type and required output.[5]
TL;DR
- Wan 3.0 accepts 2–30-second output requests; MiniMax H3 accepts 4–15 seconds. Longer continuous generation is the clearest Wan advantage.[1][2]
- reAPI lists Wan 3.0 at 480P through 4K and H3 at 768P or 2K. Wan document/webpage input stops at 1080P.[1][2]
- The displayed 2K output rate is $0.199/s for Wan and $0.119/s for H3 on the observation date. Input-related charges still matter.[3][4]
- Both support image, video and audio references. Do not add independent per-type limits together and assume every mixed combination is accepted.[1][2][5]
- MiniMax publishes H3-Base weights. Its full documented system includes additional context-processing and 2K regeneration components; a local base model is not an automatic replica of a hosted 2K service.[5]
Wan 3.0 vs MiniMax H3: the limits that affect a shot
The following table describes the public reAPI request surfaces. Model names alone are insufficient because another product can expose a different subset of the same model's capabilities.[1][2]
| Requirement | Wan 3.0 on reAPI | MiniMax H3 on reAPI |
|---|---|---|
| Output duration | 2–30 seconds | 4–15 seconds |
| Output tiers | 480P, 720P, 1080P, 2K, 4K | 768P, 2K |
| Text and frame input | Supported | Supported |
| Reference images | Up to 10 | Up to 9 |
| Reference videos | Up to 5; total input footage ≤15s | Up to 3; total input footage ≤15s |
| Reference audio | Up to 5; total ≤15s | Up to 3; total ≤15s; cannot be used alone |
| Document/public webpage input | Supported through 1080P | No corresponding field documented |
| Generated audio | On by default; switch available | Native stereo audio |
| Dedicated edit mode | Not exposed in this API | No separate edit endpoint in this schema |
| Output framing field | size | aspect_ratio, with mode-specific restrictions |
For Wan, reference-video duration plus output duration must also stay within thirty seconds. A fifteen-second reference therefore leaves at most fifteen output seconds under that rule. The headline thirty-second output ceiling is not available with every reference-video request.[1]
MiniMax's official H3 reference documentation specifies at most twelve files across mixed input types. reAPI documents the separate nine/three/three maxima, but those numbers do not establish a fifteen-file combination. Keep mixed reference sets within the stricter documented limit unless the service confirms another supported combination.[5][2]
Continuous length and native audio solve different problems
Wan is the straightforward choice when the required deliverable is a twenty-second uninterrupted scene. H3's fifteen-second ceiling cannot satisfy that duration in one request. You would need a different shot plan or an edit between clips.[1][2]
That does not make a thirty-second take the right default. If the finished piece cuts between close-ups and wide shots, separate short generations may be easier to review and replace. Decide whether the shot must remain continuous before paying for duration that editing will discard. This is an editorial decision, not a claim about either model's measured reliability.
Both models can produce sound with the picture. Wan has an audio switch, while H3's published output includes native stereo audio. A reference-audio upload guides generation; it should not be treated as a promise that an uploaded soundtrack is copied unchanged.[1][2]
For dialogue, compare the actual wording, pronunciation, timing and unwanted sound in the resulting file. A specification such as native audio confirms a capability, but it does not establish that one model will deliver the better performance for your script. Include the spoken lines and required room ambience in the evaluation prompt.
Match reference and frame controls before migrating a request
Wan supports document and public webpage input through file_url and link_url. Those fields are mutually exclusive and unavailable at 2K or 4K. Choose Wan for this documented input path when the source is a brief, deck or public page; first check that a 1080P-or-lower deliverable meets your needs.[1]
H3 has no equivalent document or webpage field in reAPI's published schema. You can prepare a prompt and supported media yourself, but that is additional work, not the same input capability.[2]
For image to video, the field names differ. Wan uses roles such as first_frame and last_frame in image_with_roles. H3 uses first_frame_url and last_frame_url; its image mode derives framing from the supplied image and rejects aspect_ratio. Copying an entire JSON request and changing only model is therefore insufficient.[1][2]
| Intended control | Wan field | H3 field |
|---|---|---|
| Explicit first frame | image_with_roles with first_frame | first_frame_url |
| Explicit last frame | image_with_roles with last_frame | last_frame_url |
| Reference image | image_with_roles with reference_image | reference_image_urls |
| Reference footage | video_urls | reference_video_urls |
| Reference sound | audio_urls | reference_audio_urls |
Both schemas separate frame input from reference input and require public HTTP(S) media URLs. Design the request around one family rather than mixing all available controls.[1][2]
Editing needs a further distinction. H3's model page describes reference-driven operations such as subject replacement and relighting, but the API exposes these through reference generation. Wan's reAPI endpoint has no dedicated editing mode even though other Wan interfaces advertise editing. If a source-video edit is mandatory, inspect the precise contract or the separate Wan 2.7 comparison before choosing.[4][1]
Compare output prices at the same tier and duration
The table below uses the named output rows from the rendered reAPI price pages on September 14, 2026. Currency is USD. Each ten-second estimate multiplies the displayed rate by ten; it assumes no reference video and no extra-image charge.[3][4]
| Model and output tier | Displayed output rate | 10-second display-rate estimate |
|---|---|---|
| Wan 3.0, 480P | $0.035/s | $0.350 |
| Wan 3.0, 720P | $0.068/s | $0.680 |
| Wan 3.0, 1080P | $0.136/s | $1.360 |
| Wan 3.0, 2K | $0.199/s | $1.990 |
| Wan 3.0, 4K | $0.229/s | $2.290 |
| MiniMax H3, 768P | $0.074/s | $0.740 |
| MiniMax H3, 2K | $0.119/s | $1.190 |
These estimates are not exact invoices. Display rates are rounded, and final credits are rounded at the task level. The pages and returned usage are the references for a new request. The shared 2K tier gives a direct listed-rate comparison; 720P and 768P are different output tiers and should not be described as identical resolution.[3][4][1][2]
The input material can change the bill. Both models add the duration of reference footage to billable work. H3 also charges for the sixth through ninth reference image; its table displayed $0.036 for each additional image. That amount is an image surcharge, not a video-output price.[2][4][1]
For example, a ten-second result guided by six seconds of reference footage represents sixteen billable seconds before any H3 image surcharge. Compare that complete job, not the output duration alone. For Wan's automatic-duration option, the documented thirty-second billing ceiling makes an explicit duration more predictable.[1][2]
Evaluate quality and waiting time with the same acceptance test
The evidence here supports capability and price decisions, not a universal quality ranking. To compare a shot both models can produce, hold the requested duration, framing, source assets, dialogue and target output tier constant. A comparison at 2K is feasible in both schemas; document input is not a shared test case.[1][2]
Before generating, define acceptance criteria. Check whether identity persists, objects remain coherent, camera motion matches the instruction, speech is correct and the file meets delivery requirements. Record submission time, completion time, output duration, charged credits and whether the clip was usable. Keep failures in the record instead of comparing only the most attractive samples.
A cheaper request can cost more per accepted shot if it requires more attempts. Document input can save preparation work even when its per-second rate is higher. Measure both when testing your own workflow.
If local deployment is a deciding factor, MiniMax's official release includes H3-Base weights and documents the additional system components. Review the H3 local-versus-API guide and license discussion. For another hosted comparison, see Seedance 2.5 vs Wan 3.0 or H3 vs Seedance 2.5.[5]
FAQ
wan 3.0 vs minimax h3
Choose Wan for 16–30-second output or document/webpage input. Choose H3 when its 4–15-second output fits and the lower listed 2K rate is useful.[1][2][3][4]
wan 3 vs minimax h3
Compare the whole request. At 2K, H3 has the lower displayed output rate on the checked date, but reference-video seconds and additional H3 images can increase the total.[3][4][2]
wan 3.0 use both video and image reference
Yes, in the reference family. Wan accepts image and video references together; keep frame roles separate and respect its reference-video/output duration limit.[1]
wan 3.0 image to video
Wan supports explicit opening and closing frames. H3 does too, but uses different field names and derives image-mode framing from the source.[1][2]
wan 3.0 price
The checked reAPI table displayed $0.035–$0.229 per output second across 480P–4K. Use the selected tier and complete billable duration; this displayed range is not a final task quote.[3][1]
wan 3.0 slow
This article does not establish a speed winner. Compare elapsed time for the same duration, tier and input type, then evaluate cost per accepted result rather than anecdotal waiting times.
Choose the model around the shot's non-negotiable requirement
For Wan 3.0 vs MiniMax H3, decide length and input format first, then compare cost. Wan fits longer continuous generation and document-driven requests. H3 is a sensible first test for shorter 2K work when its inputs fit. Use the current model pages to budget, then keep the model that meets your shot's acceptance criteria at a workable cost.
References
- reAPI. Wan 3.0 | reAPI. Retrieved September 14, 2026 from reapi.ai/docs/wan-3-0
- reAPI. MiniMax H3 API — Hailuo 03 Video, T2V, I2V & Reference | reAPI. Retrieved September 14, 2026 from reapi.ai/docs/minimax-h3
- reAPI. Wan 3.0 — 30-Second Omni-Modal Video Generation. Retrieved September 14, 2026 from reapi.ai/models/wan-3-0
- reAPI. MiniMax H3 — Multimodal Video Generation with Native Audio. Retrieved September 14, 2026 from reapi.ai/models/minimax-h3
- MiniMax. Open General Intelligence: MiniMax H3 Is Now Open Source - MiniMax News | MiniMax. Retrieved September 14, 2026 from www.minimax.io/news/minimax-h3-open-source
Author

Categories
More Posts

AI Music Video Generator API: Full Songs, Photos, and Lyrics
Turn a 10-second to 5-minute song and 1–7 reference images into a complete music video, with API examples, subtitle options, and exact cost tables.


MiniMax H3 vs Seedance 2.5: Which AI Video Model Wins?
Compare MiniMax H3 vs Seedance 2.5 on duration, 2K output, native audio, multimodal references, editing, pricing, API access, and best use cases.


Run a 70B LLM on a 4GB GPU with AirLLM: The Honest Guide
Can a 70B LLM really run on a 4GB GPU? Learn how AirLLM streams layers from disk, what hardware it still needs, how to try it, and why it is slow.
