SGLang H3 /v1/videos vs Sume /v1/videos: same path, new fields
A self-hosted MiniMax H3 SGLang server and Sume both expose /v1/videos, but the fields differ: seconds vs duration, seed accepted vs not listed. Small adapter.

The self-hosted SGLang server for MiniMax H3 and Sume both have POST /v1/videos, but they are not drop-in replacements. SGLang takes seconds, task, num_inference_steps, flow_shift and seed in the body; Sume takes duration, resolution and aspect_ratio and lists seed: false. A thin adapter keeps one client for both.
Which fields map?
| Meaning | SGLang self-host | Sume `minimax-h3` |
|---|---|---|
| Model | MiniMaxAI/MiniMax-H3 | minimax-h3 |
| Prompt | prompt | prompt |
| Length | seconds (card says 4 to 15) | duration, 5 to 15 |
| Mode | task | implied by frame_images or input_references |
| Steps and flow shift | num_inference_steps, flow_shift, audio_flow_shift | not exposed |
| Seed | seed | not listed; seed: false |
What does the adapter look like?
The function below turns an SGLang-style body into a Sume body and refuses lengths Sume does not take. Run it as is.
def to_sume(body: dict) -> dict:
seconds = int(body["seconds"])
if not 5 <= seconds <= 15:
raise ValueError("minimax-h3 on Sume takes 5 to 15 seconds")
return {
"model": "minimax-h3",
"prompt": body["prompt"],
"duration": seconds,
"resolution": "768p",
}
print(to_sume({"prompt": "A tram in fog", "seconds": 8, "seed": 1101}))Why keep a switch at all?
A self-hosted server is cheap when it is busy and costly when it is idle; a hosted API is the reverse. A client that can point at either lets you send routine load to your GPUs and send overflow, or a clip longer than your node can take quickly, to the hosted id. The switch is one function plus a base URL and a key.
Keep the request log in your own format, with the backend as a field, so you can see which route made each clip and what it cost. Sume also accepts an Idempotency-Key header, so a retried request returns the original job instead of a second charge.
What stays different after the adapter?
Polling and results. On Sume the submit response has polling_url; statuses are pending, in_progress, completed, failed and cancelled; the video is at unsigned_urls[0]. I did not read SGLang's response schema, so confirm it against your server before you share a client. The Sume-side shape is in the video docs.
The adapter drops seed on purpose. The Sume catalog lists seed: false for this model, so do not forward it.
Sources
Related posts
More in Developers
- Size a Sume submit wave: limit 100, 30 processing, 10 queued gives 60
How many new jobs can a Sume workspace take? max(0, limit - active - queued), capped by queue_capacity_remaining. Worked numbers and a TypeScript helper.
- Smoke-test 1:4, 4:1, 1:8 and 8:1 on Nano Banana 2.1 for $0.30
Four calls at 0.5K, one per new ratio, cost $0.30 on Sume ($0.80 at 4K). A script that prints status and cost per ratio, and what a 400 or 202 means.
- Sora, Veo and Omni preview shutdown dates in one table (Oct 2026)
OpenAI removed the Sora Videos API on Sept 24, 2026. Google ends Veo 3.1 and Omni preview ids Oct 22. One dated table, and how to read Sume's live list.
- Make an SRT file from Sume STT word timings in Python (7 words a cue)
Sume STT always returns words[] with word, start and end in seconds. This Python function turns the list into an SRT file, seven words a cue.
Written by Sume