Self-host MiniMax H3 or use an API: a decision table
MiniMax released H3 as an open-weight omni-modal model on Jul 31, 2026. A decision table for running it yourself versus calling a hosted video API.

Run MiniMax H3 yourself if you need control over weights, LoRAs or data location and can operate GPUs. Call a hosted API if you want a finished MP4 per request, billing and retries handled for you. The MiniMax model page, read 2026-10-03, says H3 was released on Jul 31, 2026 as an open-weight omni-modal model.
What the vendor page and search suggest
The MiniMax page confirms the open-weight release. Search autocomplete for the model that day suggested four follow-ups: Hugging Face, ComfyUI, LoRA and local. That is a hint about what people want to do with the weights, not a measure of how many succeed.
The decision table
The right column describes how Sume documents its own hosted route. It is not a claim about what any particular local setup costs.
| Question | Self-host the weights | Hosted API |
|---|---|---|
| Who provisions GPUs | You | Provider, you pay per output second |
| Custom LoRA or fine-tune | Possible, you carry the training | Only the fields the API exposes |
| Request shape | Whatever your pipeline defines | A documented async job |
| Retries and idempotency | You build them | On Sume, an Idempotency-Key makes a replay return the original job |
| Billing | Your infrastructure bill | On Sume, USD balance reserved at submit, provider list x 1.25 |
| Upgrades | You pull and re-test | Catalog changes ship server side |
How a hosted job behaves
On Sume the call is asynchronous. You submit to POST /v1/videos, receive a job id and polling URL, then poll GET /v1/videos/{jobId} until completed and download from unsigned_urls. The jobs guide says to use exponential backoff and never resubmit a paid request just because a local process timed out.
The catalog entry minimax-h3 accepts 5-15 seconds at native 480p or 768p. If your self-hosted build supports a different range, that is a difference you own.
Questions to settle first
Before you commit either way, answer four questions in writing. How many seconds of video do you need per month, and how spiky is that demand? Which team owns GPU operations, and what happens at 3 a.m. when a node fails? Does any input data have to stay on your own hardware? And which output limits do you need, such as duration and resolution?
Hosted catalogs state their limits per model, so you can compare them to your answers directly. For self-hosting, you will establish the same limits by test.
A reasonable split
Many teams prototype against an API to settle prompts and durations, then decide whether volume justifies self-hosting. Moving later costs you a client rewrite, not a prompt rewrite, if you keep your own request model thin. Keep a record of the Sume job_id next to each asset so you can compare outputs from both paths.
Sources
Related posts
More in Models
- Veo 3.1 extend adds 7 s up to 20 times, 720p only: plan a long clip
Google's Veo docs let you extend a clip by 7 seconds up to 20 times, at 720p only. Read the arithmetic, then compare it with one-call 30 second models on Sume.
- Voxtral Mini Transcribe 2 and Realtime v26.02: what Mistral lists
Mistral lists Voxtral Mini Transcribe 2, Voxtral Realtime v26.02 and Voxtral TTS v26.03. How to prepare video audio for any transcription model.
- Wan 3.0 doubles clip length to 30 seconds: Alibaba ids vs Sume
Alibaba Model Studio lists wan3.0-video and wan3.0-video-prime at 2 to 30 seconds, up from 15 on Wan 2.7. What Sume's wan-3.0 accepts and a 30 s request.
- What is FLUX 3? The family map: image, video, audio, action
BFL's docs describe FLUX 3 as one family covering image, video with synchronized audio, audio and action. Which pieces have open weights and which do not.
Written by Sume