One request for four Sume video models: 1080p, 5 to 10 s, 16:9 or 9:16

Kling 3, Wan 3.0, H3 Max and Omni share one request shape: 1080p, 5 to 10 seconds, 16:9 or 9:16. Compute the intersection in Python and price each model.

4 min readSume
All posts

Four Video Router models accept one identical request: resolution: "1080p", a duration from 5 to 10 seconds, and aspect_ratio of 16:9 or 9:16. Those are kling-3, wan-3.0, minimax-h3-max and gemini-omni-flash-1.1. Anywhere outside that envelope at least one of them says no, and that is useful when you want to fan a single prompt out to several models and compare the clips. The envelope below is computed from the catalog described in the Video Router docs.

What does each model allow?

Each row is copied from the catalog's capabilities for that id.

Resolution, duration and aspect ratios per model, Sume Video Router catalog, read 2026-10-03
Model idResolutionsDurationAspect ratios
kling-3720p, 1080p4-15 s16:9, 9:16, 1:1
wan-3.0480p, 720p, 1080p2-30 sauto, adaptive, 16:9, 4:3, 1:1, 3:4, 9:16
minimax-h3-max480p, 768p, 1080p5-15 sadaptive, 21:9, 16:9, 4:3, 1:1, 3:4, 9:16
gemini-omni-flash-1.1360p, 720p, 1080p, 4K3-10 s16:9, 9:16

How do I compute the shared range?

Resolution and ratio are set intersections. Duration is the largest minimum to the smallest maximum: Kling starts at 4 and H3 Max at 5, so the floor is 5, and Omni stops at 10, so the ceiling is 10. The script below does the arithmetic and prints the result.

LIMITS = {
    "kling-3": ({"720p", "1080p"}, (4, 15), {"16:9", "9:16", "1:1"}),
    "wan-3.0": ({"480p", "720p", "1080p"}, (2, 30), {"16:9", "4:3", "1:1", "3:4", "9:16"}),
    "minimax-h3-max": ({"480p", "768p", "1080p"}, (5, 15),
                      {"21:9", "16:9", "4:3", "1:1", "3:4", "9:16"}),
    "gemini-omni-flash-1.1": ({"360p", "720p", "1080p", "4K"}, (3, 10), {"16:9", "9:16"}),
}
rows = list(LIMITS.values())
res = set.intersection(*(r[0] for r in rows))
ratios = set.intersection(*(r[2] for r in rows))
lo = max(r[1][0] for r in rows)
hi = min(r[1][1] for r in rows)
print(sorted(res), sorted(ratios), lo, hi)

What does the same clip cost on each?

Same request, different bills. A 10-second 1080p clip costs the amounts below, each at provider list times 1.25 rounded up per clip.

10-second 1080p clip, Sume price = list x 1.25 rounded up (Sume Video Router catalog, read 2026-10-03)
Model idList per secondSume price for 10 s
kling-3 (audio on)$0.168$2.10
wan-3.0$0.200$2.50
minimax-h3-max$0.160$2.00
gemini-omni-flash-1.1$0.150$1.88

Why does fan-out help?

Comparing models on the same prompt is the cheapest way to learn which one fits your brand's look, and a shared body removes the variable you do not want to test. Run the four requests with the same prompt and a distinct idempotency key for each, then compare the clips side by side.

Keep the duration at 5 seconds for the first pass. A 5-second 1080p clip costs $1.05 on Kling 3 with audio on, $1.25 on Wan 3.0, $1.00 on H3 Max and $0.94 on Omni, so the whole comparison is $4.24 before you pick a winner for the longer run.

How do I send the same body to all four?

The shared body has five fields: model, prompt, resolution, duration and aspect_ratio. Only model changes between requests. The Video Router docs show the pattern with POST /v1/video-router/generate, an Idempotency-Key header and mode: "async", and the response is the usual async job envelope that you poll for the finished video.

Send the four requests in parallel, each with its own idempotency key, for example compare-kling-001, compare-wan-001, compare-h3max-001 and compare-omni-001. A replay with the same key returns the original job instead of creating a second one, so a client retry after a network error cannot double the bill. Then poll each job and download the results when they finish.

If a model fails, the failure is per job: the other three are unaffected. That is the main reason to fan out with separate jobs rather than hoping a single call covers every model.

When the comparison is done, move to the narrower envelope of the winner. Wan 3.0 is the one to pick if you want to go longer than 10 seconds or add references; Kling 3 is the one to pick if the budget is the constraint; H3 Max is the pick for 21:9 or for many references; Omni is the pick for 4K. Each of those facts is a catalog field, not a taste judgment.

What should I leave out of a shared body?

Stay with plain fields. Do not send generate_audio to H3 Max or Omni, because both always produce sound and the Omni constraint says generate_audio: false is rejected. Do not send reference media, because Kling takes none. Do not send size, seed or provider.options, which every model refuses. An idempotency key per model keeps retries from double-booking, since a replay returns the original job.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume