
MiniMax H3 Local vs API: VRAM, Speed, License, and Cost
What do official MiniMax and ComfyUI docs establish about local H3 hardware? Compare the documented workflow with hosted 768P and 2K API costs.
Run MiniMax H3 locally when the official optimized workflow fits your hardware and you need repeated iteration or private inputs; use an API when you need occasional finished clips, hosted 2K output, concurrency, or no setup work. The official ComfyUI guide says dynamic offloading can launch the workflow on an RTX 3060, but launch alone is not a speed benchmark or proof that different VRAM tiers deliver equivalent results.[1]
The phrase “runs locally” needs three numbers attached to it: resolution, duration and runtime. Without them, a low-memory launch and a server deployment appear to be the same result.
What local H3 currently means
| Setup | Officially documented result | Important qualification |
|---|---|---|
| Official optimized ComfyUI package | 42.5GB total model footprint; dynamic offload can launch on RTX 3060 | Launch is not a speed benchmark |
| Official SGLang example | Full deployment across four GPUs | Server route, not a consumer single-card recipe |
ComfyUI's official H3 guide sets the native 16:9 canvas at 1344×768 and packages pruned INT8 weights, an NVFP4 text encoder, separate video and audio VAEs and optional Turbo LoRAs.[1] Its published optimization reduced the full-precision footprint from 123.6GB to 42.5GB. Offloading lets a smaller card start the workflow by moving parts through system memory; it does not make those parts disappear.
That is why VRAM capacity alone does not answer whether a local workflow is useful. Resolution, duration, runtime, and the amount of offloading all belong in the same test result.
What the official docs establish about local hardware
The cited official material establishes two concrete deployment paths: an optimized ComfyUI workflow that uses dynamic offloading, and a multi-GPU SGLang example. It does not publish a comparable performance grid for 8GB, 12GB, and 24GB consumer systems.
Before treating any local result as a purchasing recommendation, ask three direct questions:
- Can the exact workflow load without an out-of-memory error?
- What resolution and duration does it produce, and how long does one attempt take?
- Does the requested output use only the native local workflow, or an official hosted context or regeneration service?
MiniMax's model card distinguishes the local native workflow from the full 2K path. The published 2K workflow uses local deployment together with official Context-IR or regeneration APIs.[2] Local 768p and hosted 2K are therefore different output products, not two prices for an identical job.
MiniMax H3 API cost
Prices checked August 23, 2026:
| Hosted route | Per second | 5 seconds | 15 seconds |
|---|---|---|---|
| MiniMax direct, 768P | $0.080 | $0.400 | $1.200 |
| reAPI, 768P | $0.074 | $0.370 | $1.110 |
| MiniMax direct, 2K | $0.130 | $0.650 | $1.950 |
| reAPI, 2K | $0.119 | $0.595 | $1.785 |
MiniMax's official price sheet charges by output duration. Input-video duration is also billed at the selected resolution, the first five reference images are free, later images add a charge and audio references are free.[3]
The reAPI MiniMax H3 model page publishes separate 768P and 2K rates.[4] These are dated numbers, not permanent promises; use the live page for a budget.
Hardware break-even is measured in attempts
Suppose H3 requires a $500 hardware upgrade. Ignore electricity and setup time first.
- At $1.11 for a 15-second 768P API clip, $500 buys about 450 attempts.
- At $1.785 for a 15-second 2K clip, it buys about 280 attempts.
- A $1,000 machine budget moves those figures to about 900 attempts at 768P or 560 at 2K.
The correct unit is attempts, not delivered videos. If one accepted shot takes five generations, 450 attempts may become 90 finished shots. Local generation benefits most when rejected drafts dominate the workload.
The arithmetic also changes when the GPU already sits on the desk. Its purchase cost is sunk, and local H3 becomes attractive quickly for unlimited low-resolution iteration, offline reference material, custom nodes and full workflow control.
API access earns its cost when volume is irregular, a 2K deliverable is required, several jobs should run without occupying a workstation or the team would rather spend $0.37 testing a prompt than spend a weekend tuning memory settings.
Check the license before buying the GPU
MiniMax H3 is open-weight, not Apache- or MIT-licensed. The current community license excludes the United States, European Union, United Kingdom and South Korea from its applicable territory and directs users there to obtain separate written authorization.[5]
Public download and permission for a deployment are separate questions. The license also contains commercial revenue and attribution conditions. Our MiniMax H3 license guide explains the clauses without trying to provide legal advice.
That distinction matters most for self-hosting because the team is taking and operating the weights directly. A hosted API is a different contractual route and should be evaluated under its service terms.
Frequently asked questions
Can MiniMax H3 run on 8GB VRAM?
The cited official materials do not establish an 8GB result. The official ComfyUI guide documents dynamic-offload launch on an RTX 3060, but that does not establish performance for an 8GB configuration.
Is 12GB VRAM enough for H3?
The official evidence cited here is insufficient to call every 12GB setup suitable. It demonstrates a specific optimized, offloaded path rather than a general 12GB performance guarantee.
Is local H3 free?
There is no per-generation API bill, but hardware, power, storage, setup time and failed local attempts still have costs. The community license can also determine whether the intended deployment is allowed.
When is the API cheaper than local H3?
For a new $500 hardware expense, the current break-even is roughly 450 fifteen-second 768P attempts or 280 fifteen-second 2K attempts on reAPI, before electricity and setup time.
Does local H3 produce the same 2K output as the hosted API?
Not through the basic native ComfyUI workflow. MiniMax describes the full 2K workflow as a combination of local deployment and official context or regeneration services.
References
- ComfyUI. MiniMax H3 official tutorial. Retrieved August 23, 2026. docs.comfy.org
- MiniMaxAI. MiniMax H3 model card and deployment examples. Retrieved August 23, 2026. huggingface.co
- MiniMax API. Pay-as-you-go pricing. Retrieved August 23, 2026. platform.minimax.io
- reAPI. MiniMax H3 model page and live pricing. Retrieved August 23, 2026. reapi.ai
- MiniMaxAI. MiniMax H3 Community License. Retrieved August 23, 2026. huggingface.co
Author

Categories
More Posts

GPT Image 2.5 Model Not Found: Fix Codex and API Access
Fix GPT Image 2.5 model not found errors by checking model IDs, Image API versus Responses fields, Codex tools and account access before retrying requests.


LLM Max Output Tokens: API Limits and Context Budgets
Compare LLM output limits, reasoning budgets, and API parameters. Separate context windows, request caps, defaults, and Claude's batch-only output extension.


Opus 5.5 Video Generation: Directing Seedance 2.5 Shots
Opus 5.5 video generation, done right: Claude writes the shot plan, GPT Image 2.5 locks the character, Seedance 2.5 renders it. Prompts, code and real costs.
