One request for four Sume video models: 1080p, 5 to 10 s, 16:9 or 9:16
Kling 3, Wan 3.0, H3 Max and Omni share one request shape: 1080p, 5 to 10 seconds, 16:9 or 9:16. Compute the intersection in Python and price each model.

Four Video Router models accept one identical request: resolution: "1080p", a duration from 5 to 10 seconds, and aspect_ratio of 16:9 or 9:16. Those are kling-3, wan-3.0, minimax-h3-max and gemini-omni-flash-1.1. Anywhere outside that envelope at least one of them says no, and that is useful when you want to fan a single prompt out to several models and compare the clips. The envelope below is computed from the catalog described in the Video Router docs.
What does each model allow?
Each row is copied from the catalog's capabilities for that id.
| Model id | Resolutions | Duration | Aspect ratios |
|---|---|---|---|
| kling-3 | 720p, 1080p | 4-15 s | 16:9, 9:16, 1:1 |
| wan-3.0 | 480p, 720p, 1080p | 2-30 s | auto, adaptive, 16:9, 4:3, 1:1, 3:4, 9:16 |
| minimax-h3-max | 480p, 768p, 1080p | 5-15 s | adaptive, 21:9, 16:9, 4:3, 1:1, 3:4, 9:16 |
| gemini-omni-flash-1.1 | 360p, 720p, 1080p, 4K | 3-10 s | 16:9, 9:16 |
How do I compute the shared range?
Resolution and ratio are set intersections. Duration is the largest minimum to the smallest maximum: Kling starts at 4 and H3 Max at 5, so the floor is 5, and Omni stops at 10, so the ceiling is 10. The script below does the arithmetic and prints the result.
LIMITS = {
"kling-3": ({"720p", "1080p"}, (4, 15), {"16:9", "9:16", "1:1"}),
"wan-3.0": ({"480p", "720p", "1080p"}, (2, 30), {"16:9", "4:3", "1:1", "3:4", "9:16"}),
"minimax-h3-max": ({"480p", "768p", "1080p"}, (5, 15),
{"21:9", "16:9", "4:3", "1:1", "3:4", "9:16"}),
"gemini-omni-flash-1.1": ({"360p", "720p", "1080p", "4K"}, (3, 10), {"16:9", "9:16"}),
}
rows = list(LIMITS.values())
res = set.intersection(*(r[0] for r in rows))
ratios = set.intersection(*(r[2] for r in rows))
lo = max(r[1][0] for r in rows)
hi = min(r[1][1] for r in rows)
print(sorted(res), sorted(ratios), lo, hi)What does the same clip cost on each?
Same request, different bills. A 10-second 1080p clip costs the amounts below, each at provider list times 1.25 rounded up per clip.
| Model id | List per second | Sume price for 10 s |
|---|---|---|
| kling-3 (audio on) | $0.168 | $2.10 |
| wan-3.0 | $0.200 | $2.50 |
| minimax-h3-max | $0.160 | $2.00 |
| gemini-omni-flash-1.1 | $0.150 | $1.88 |
Why does fan-out help?
Comparing models on the same prompt is the cheapest way to learn which one fits your brand's look, and a shared body removes the variable you do not want to test. Run the four requests with the same prompt and a distinct idempotency key for each, then compare the clips side by side.
Keep the duration at 5 seconds for the first pass. A 5-second 1080p clip costs $1.05 on Kling 3 with audio on, $1.25 on Wan 3.0, $1.00 on H3 Max and $0.94 on Omni, so the whole comparison is $4.24 before you pick a winner for the longer run.
How do I send the same body to all four?
The shared body has five fields: model, prompt, resolution, duration and aspect_ratio. Only model changes between requests. The Video Router docs show the pattern with POST /v1/video-router/generate, an Idempotency-Key header and mode: "async", and the response is the usual async job envelope that you poll for the finished video.
Send the four requests in parallel, each with its own idempotency key, for example compare-kling-001, compare-wan-001, compare-h3max-001 and compare-omni-001. A replay with the same key returns the original job instead of creating a second one, so a client retry after a network error cannot double the bill. Then poll each job and download the results when they finish.
If a model fails, the failure is per job: the other three are unaffected. That is the main reason to fan out with separate jobs rather than hoping a single call covers every model.
When the comparison is done, move to the narrower envelope of the winner. Wan 3.0 is the one to pick if you want to go longer than 10 seconds or add references; Kling 3 is the one to pick if the budget is the constraint; H3 Max is the pick for 21:9 or for many references; Omni is the pick for 4K. Each of those facts is a catalog field, not a taste judgment.
What should I leave out of a shared body?
Stay with plain fields. Do not send generate_audio to H3 Max or Omni, because both always produce sound and the Omni constraint says generate_audio: false is rejected. Do not send reference media, because Kling takes none. Do not send size, seed or provider.options, which every model refuses. An idempotency key per model keeps retries from double-booking, since a replay returns the original job.
Sources
Related posts
More in Developers
- OpenAI Agents SDK client_session_timeout_seconds with Sume jobs_wait
In the OpenAI Agents SDK for Python, client_session_timeout_seconds sets the MCP read timeout. Set it above Sume's 55-second jobs_wait cap, or 0 to disable it.
- OpenAI Agents SDK max_retry_attempts and Sume paid calls
The OpenAI Agents SDK can retry MCP list_tools and call_tool. A retried Sume create must repeat its idempotency_key or you pay twice; set max_spend_usd too.
- Move OpenAI GPT Image 2.5 calls to Sume: field-by-field mapping
Which OpenAI gpt-image-2.5 parameters carry over to Sume's POST /v1/images, which change name, and which return 400 unsupported_parameter.
- OpenAI transcription 25 MB limit: how many minutes of wav fit?
OpenAI caps transcription uploads at 25 MB. A 16 kHz mono wav fills that in about 13 minutes, so detach long videos to mp3 or ranges first.
Written by Sume