
MiniMax H3 Local vs API: VRAM, Speed, License, and Cost
Can MiniMax H3 run on 8GB, 12GB, or 24GB VRAM? Compare public local results with hosted 768P and 2K API costs and the H3 license limits.
Run MiniMax H3 locally when you already own a 24GB-class GPU and need high iteration volume or private inputs; use an API when you need occasional finished clips, hosted 2K output, concurrency or no setup work. H3 can launch on 8GB hardware, but the public 8GB result is a five-second 608×352 clip, not a production-equivalent replacement for hosted 768P or 2K generation.[2]
The phrase “runs locally” needs three numbers attached to it: resolution, duration and runtime. Without them, an 8GB experiment and a server deployment appear to be the same result.
What local H3 currently means
| Setup | Published result | Important qualification |
|---|---|---|
| Official optimized ComfyUI package | 42.5GB total model footprint; dynamic offload can launch on RTX 3060 | Launch is not a speed benchmark |
| 8GB VRAM + 16GB RAM | 608×352, 5 seconds in 1m35s | Experimental quantized workflow |
| RTX 3090 24GB + 32GB RAM | 832×480, 15 seconds with stereo audio | Required 40GB swap |
| Official SGLang example | Full deployment across four GPUs | Server route, not a consumer single-card recipe |
ComfyUI's official H3 guide sets the native 16:9 canvas at 1344×768 and packages pruned INT8 weights, an NVFP4 text encoder, separate video and audio VAEs and optional Turbo LoRAs.[1] Its published optimization reduced the full-precision footprint from 123.6GB to 42.5GB. Offloading lets a smaller card start the workflow by moving parts through system memory; it does not make those parts disappear.
That is why VRAM alone gives the wrong answer. System RAM, swap, storage speed, target megapixels and how often the workflow moves tensors across the bus all affect whether “works” feels usable.
The public 8GB workflow used low-VRAM flags, a smaller text encoder and a 608×352 target.[2] A separate RTX 3090 guide identified 32GB of system RAM as the binding constraint and expanded swap to 40GB.[3]
An 8GB yes is a preview tier, not a production recommendation
The current evidence supports a more useful hardware split:
- 8–12GB VRAM: experiments, reduced resolution and learning the workflow. Always attach a runtime and output size to any success claim.
- 24GB VRAM: the first consumer tier that looks practical for repeated local use, provided the machine also has enough RAM and fast storage.
- Multi-GPU server: the official base deployment path when throughput and full model components matter more than workstation simplicity.
MiniMax's model card also distinguishes the local native workflow from the full 2K path. The published 2K workflow uses local deployment together with official Context-IR or regeneration APIs.[4] Local 768p and hosted 2K are therefore different output products, not two prices for an identical job.
MiniMax H3 API cost
Prices checked August 23, 2026:
| Hosted route | Per second | 5 seconds | 15 seconds |
|---|---|---|---|
| MiniMax direct, 768P | $0.080 | $0.400 | $1.200 |
| reAPI, 768P | $0.074 | $0.370 | $1.110 |
| MiniMax direct, 2K | $0.130 | $0.650 | $1.950 |
| reAPI, 2K | $0.119 | $0.595 | $1.785 |
MiniMax's official price sheet charges by output duration. Input-video duration is also billed at the selected resolution, the first five reference images are free, later images add a charge and audio references are free.[5]
The reAPI MiniMax H3 model page publishes separate 768P and 2K rates.[6] These are dated numbers, not permanent promises; use the live page for a budget.
Hardware break-even is measured in attempts
Suppose H3 requires a $500 hardware upgrade. Ignore electricity and setup time first.
- At $1.11 for a 15-second 768P API clip, $500 buys about 450 attempts.
- At $1.785 for a 15-second 2K clip, it buys about 280 attempts.
- A $1,000 machine budget moves those figures to about 900 attempts at 768P or 560 at 2K.
The correct unit is attempts, not delivered videos. If one accepted shot takes five generations, 450 attempts may become 90 finished shots. Local generation benefits most when rejected drafts dominate the workload.
The arithmetic also changes when the GPU already sits on the desk. Its purchase cost is sunk, and local H3 becomes attractive quickly for unlimited low-resolution iteration, offline reference material, custom nodes and full workflow control.
API access earns its cost when volume is irregular, a 2K deliverable is required, several jobs should run without occupying a workstation or the team would rather spend $0.37 testing a prompt than spend a weekend tuning memory settings.
Check the license before buying the GPU
MiniMax H3 is open-weight, not Apache- or MIT-licensed. The current community license excludes the United States, European Union, United Kingdom and South Korea from its applicable territory and directs users there to obtain separate written authorization.[7]
Public download and permission for a deployment are separate questions. The license also contains commercial revenue and attribution conditions. Our MiniMax H3 license guide explains the clauses without trying to provide legal advice.
That distinction matters most for self-hosting because the team is taking and operating the weights directly. A hosted API is a different contractual route and should be evaluated under its service terms.
Frequently asked questions
Can MiniMax H3 run on 8GB VRAM?
Yes, in an experimental low-memory workflow. The public result cited here generated five seconds at 608×352 in 1 minute 35 seconds. It does not establish native 768P or 2K production performance.
Is 12GB VRAM enough for H3?
It is enough for experimentation with quantization and offloading. Published evidence does not support treating 12GB as equivalent to a comfortable 24GB workflow; system RAM and target resolution remain significant.
Is local H3 free?
There is no per-generation API bill, but hardware, power, storage, setup time and failed local attempts still have costs. The community license can also determine whether the intended deployment is allowed.
When is the API cheaper than local H3?
For a new $500 hardware expense, the current break-even is roughly 450 fifteen-second 768P attempts or 280 fifteen-second 2K attempts on reAPI, before electricity and setup time.
Does local H3 produce the same 2K output as the hosted API?
Not through the basic native ComfyUI workflow. MiniMax describes the full 2K workflow as a combination of local deployment and official context or regeneration services.
References
- ComfyUI. MiniMax H3 official tutorial. Retrieved August 23, 2026. docs.comfy.org
- r/StableDiffusion. MiniMax H3 on 8GB VRAM and 16GB RAM. Retrieved August 23, 2026. reddit.com
- tonyd2wild. MiniMax H3 local RTX 3090 ComfyUI guide. Retrieved August 23, 2026. github.com
- MiniMaxAI. MiniMax H3 model card and deployment examples. Retrieved August 23, 2026. huggingface.co
- MiniMax API. Pay-as-you-go pricing. Retrieved August 23, 2026. platform.minimax.io
- reAPI. MiniMax H3 model page and live pricing. Retrieved August 23, 2026. reapi.ai
- MiniMaxAI. MiniMax H3 Community License. Retrieved August 23, 2026. huggingface.co
Author

Categories
More Posts

FLUX 3 vs Seedance 2.5: Keyframes, Duration, and Cost
Choose FLUX 3 or Seedance 2.5 by maximum duration, ordered keyframes, references, editing, native audio, Draft workflow, resolution, and API cost.


GPT Image 2 + Seedance 2.0: A Character Consistency Workflow
Use GPT Image 2 and Seedance 2.0 to reduce character drift across AI video shots. Build identity references, animate controlled frames, and chain clips.


Cheapest Veo 3.1 API in 2026: Every Provider's Real Price
Veo 3.1 API prices run from $0.40/sec on Google direct to $0.046 per 8-second clip on reAPI. Full price comparison across five providers, May 2026.
