Test one video prompt on ten models for under $7 (Python sweep)
A 5-second, lowest-resolution sweep of one prompt over ten Sume video ids costs about $5.60 in total. The price of each row and a script that submits them.

The cheapest way to compare one prompt across Sume's video models is a 5-second sweep at each model's lowest sensible resolution: ten ids cost $5.60 in total, and $3.21 without Kling 3 and Seedance 2.5. Five seconds is inside every model's duration range, so one duration value works for the whole loop.
Cost of the sweep
Kling 3 has no resolution input here and is priced with its default sound on; Omni Flash always includes audio.
| Model id | Setting | 5-second price |
|---|---|---|
| grok-imagine-video-1.5 | 480p | $0.0625 |
| gemini-omni-flash-1.1 | 360p | $0.1875 |
| wan-3.0 | 480p | $0.3125 |
| minimax-h3 | 480p | $0.3125 |
| minimax-h3-max | 480p | $0.3125 |
| seedance-2-mini | 480p | $0.4394 |
| seedance-2-fast | 480p | $0.7031 |
| seedance-2 | 480p | $0.8789 |
| seedance-2.5 | 480p | $1.34 |
| kling-3 | default | $1.05 |
| Total | $5.60 |
The script
Each job gets its own Idempotency-Key, so rerunning the script after a timeout does not double-charge. Aspect ratio is left out so it works for the models that do not take one.
import os, requests
H = {"Authorization": "Bearer " + os.environ["SUME_API_KEY"], "Content-Type": "application/json"}
PROMPT = "A barista pours latte art, slow push-in, warm window light"
RUNS = [("grok-imagine-video-1.5", "480p"), ("gemini-omni-flash-1.1", "360p"),
("wan-3.0", "480p"), ("minimax-h3", "480p"), ("minimax-h3-max", "480p"),
("seedance-2-mini", "480p"), ("seedance-2-fast", "480p"),
("seedance-2", "480p"), ("seedance-2.5", "480p"), ("kling-3", None)]
for model, res in RUNS:
body = {"model": model, "prompt": PROMPT, "duration": 5}
if res:
body["resolution"] = res
r = requests.post("https://api.sume.com/v1/videos",
headers={**H, "Idempotency-Key": "sweep-1-" + model},
json=body, timeout=60)
print(model, r.status_code, r.json().get("polling_url") or r.json())After the sweep
Drop the two or three models whose clips do not suit the brief, keep the rest, and repeat with a second prompt. A 402 on any row means the balance did not cover that hold; the other rows are unaffected.
Keeping the sweep honest
Keep the prompt, duration and aspect handling fixed, and write down the date. Lowest-resolution clips tell you about motion, framing and prompt adherence; they do not tell you how a model looks at 1080p. Once two or three ids survive, repeat the winners at the resolution you will publish.
Sume reserves the full price at submit, so the number in the table is also the amount that must be free in the balance when you send the request. If it is not, the call fails with 402 insufficient_credits before any provider work starts.
Sources
Related posts
More in Developers
- TTS word timestamps: timestamps.words and sentence segmentation
Sume TTS accepts timestamps.words and segmentation.mode sentence so a generated voiceover can drive caption timing. Request fields, rules and a working call.
- Turn a roleplay debrief into an avatar feedback clip in Python
Take the written debrief from a roleplay or survey session and render it as a 16:9 Sume avatar clip with a retry-safe key, a 12 to 168 word check and polling.
- 12 Wan 3.0 clips in parallel in Python: ThreadPoolExecutor, width 4
A Python batch for Sume: ThreadPoolExecutor at width 4 (Pro concurrency), one Idempotency-Key per item, polling by next_poll_after_seconds. Cost included.
- Vercel 800 s max duration: do you still need a Sume webhook?
Vercel Pro allows 800 s functions and a 30-minute beta. A Sume video job can still outlast one request, so use async or webhook mode and return in seconds.
Written by Sume