A/B test two AI video models by API with a stable hash bucket
Split production video jobs between two Sume model ids by hashing a job key, so retries keep the same arm. Python code, a duration table and a 1,000-key check.

To A/B test two AI video models through one API, hash a stable job key into a bucket from 0 to 99, map the bucket to a model id, and put that model id into the Idempotency-Key. On Sume both arms are the same POST /v1/videos call with a different model string, so the only code you own is the split. The split has to be deterministic: a retried request must land in the same arm, or you pay for the same clip twice and your comparison mixes the arms.
The reason to do this now is that the OpenAI Sora API ended on 2026-09-24, according to the Magic Hour release tracker (read 2026-10-06), so many teams are re-picking a model rather than swapping one endpoint. The same tracker lists Seedance 2.5 (2026-07-31) and Kling VIDEO 3.0 (2026-02-05) as current options. Sume's catalog ships both as seedance-2.5 and kling-3; the full id list is in Sume video models.
Why must the split be derived from the job key?
A random coin flip on each submit is the common mistake. It breaks the moment a worker retries after a timeout, because the second attempt can draw the other arm. Derive the arm from something that never changes for the logical job, such as the order id or the shot id.
Hash it with SHA-256, take four bytes, and reduce modulo 100. Python's built-in hash() is salted per process, so it is the wrong tool here; a restart would reshuffle every job. Keep the shares in one list so moving from 50/50 to 90/10 is a one-line change.
What does the code look like?
Sume accepts an Idempotency-Key header on submits. A retry with the same key and the same payload returns the original job, and the same key with a different payload returns 409 idempotency_conflict, per Sume generation admission. Append the model id to the key: if you later change the shares, an old key keeps its old arm instead of conflicting.
import hashlib
ARMS = [("seedance-2.5", 50), ("kling-3", 50)]
def bucket(job_key: str) -> int:
digest = hashlib.sha256(job_key.encode()).digest()
return int.from_bytes(digest[:4], "big") % 100
def pick_model(job_key: str) -> str:
n, edge = bucket(job_key), 0
for model, share in ARMS:
edge += share
if n < edge:
return model
return ARMS[-1][0]
def build(job_key: str, prompt: str) -> dict:
model = pick_model(job_key)
return {"headers": {"Idempotency-Key": f"{job_key}:{model}"},
"body": {"model": model, "prompt": prompt, "duration": 8}}
if __name__ == "__main__":
first = build("order-1001", "A kettle boiling, macro shot")
assert build("order-1001", "A kettle boiling, macro shot") == first
counts = {m: 0 for m, _ in ARMS}
for i in range(1000):
counts[pick_model(f"order-{i}")] += 1
print(first["body"]["model"], counts)
assert all(380 < c < 620 for c in counts.values())Which parameters can both arms share?
Both arms must be legal for the same request, or one arm fails on validation and your result is about error rates, not quality. Pick a duration inside both ranges. GET /v1/videos/models returns the live catalog, and a request the catalog rejects comes back as 400 unsupported_capability, so check the pair once at deploy time.
| Model id | Duration range | Note |
|---|---|---|
| seedance-2.5 | 4 to 30 s | Tracker lists up to 30 s |
| kling-3 | up to 15 s | Tracker lists 15 s with native audio |
| wan-3.0 | 2 to 30 s | Widest range of the five |
| minimax-h3 | 5 to 15 s | 480p or 768p, not 720p |
| gemini-omni-flash-1.1 | 3 to 10 s | Default 8 s |
What should stay identical across arms?
Keep everything except model identical between arms: the prompt, the duration, the aspect ratio, and the input images. Sume rejects seed on /v1/videos with 400 unsupported_parameter, and no catalog model accepts it, so you cannot pin the randomness to compare two models on a matched pair. Compare distributions over many jobs instead of pairs of clips, and let a reviewer score each clip without seeing which arm made it.
Do not use sume/auto as an arm. It is a router, not a model: Sume documents its resolution as a pure function of the normalized request and the catalog version, never discloses which family served a request, and echoes sume/auto as the job's model. An arm you cannot name is not a comparison, so pin explicit catalog ids on both sides.
What should you record?
Log the model id, the job id, and your quality score together, and compare the arms only after each has a few hundred completed jobs. Sume does not publish a quality benchmark, and this post does not invent one: your own rubric is the measurement. Cost differs by model, so record the billed amount per job from your usage data rather than from list prices.
Ramp carefully. Start at 10 percent on the new arm, because a new model id can fail validation in ways the old one did not, and a bad arm at 50 percent doubles the blast radius. When the shares change, existing keys keep their arm, since the key embeds the model, so jobs already in flight are unaffected.
If one arm starts returning 429 queue_full, that is capacity, not quality. Drop it from the comparison for that window instead of counting the failures against the model.
Sources
Related posts
More in Developers
- arq worker that polls an AI video job with defer_by in Python
An arq task reads Sume's job status once and enqueues itself again with _defer_by from next_poll_after_seconds, giving asyncio polling without a sleep loop.
- asyncio Semaphore size for Sume image batches: accepted capacity
Size the semaphore to what Sume accepts, concurrency plus queue: Free 6, Pro 24, Startup 48, Scale 120. A fake-submit test proves the peak never exceeds it.
- asyncio TaskGroup cancels siblings: poll many Sume jobs safely
A TaskGroup cancels every other poller when one raises. For a batch of Sume video jobs that abandons waits, not jobs. Catch inside the task and return results.
- Check balance before bulk transcription: 402 and admission preview
A 402 insufficient_credits arrives before any provider work. Read GET /v1/balance and POST /v1/generation/admission-preview first and size the run.
Written by Sume