A/B test two AI video models by API with a stable hash bucket

Split production video jobs between two Sume model ids by hashing a job key, so retries keep the same arm. Python code, a duration table and a 1,000-key check.

5 min readSume
All posts

To A/B test two AI video models through one API, hash a stable job key into a bucket from 0 to 99, map the bucket to a model id, and put that model id into the Idempotency-Key. On Sume both arms are the same POST /v1/videos call with a different model string, so the only code you own is the split. The split has to be deterministic: a retried request must land in the same arm, or you pay for the same clip twice and your comparison mixes the arms.

The reason to do this now is that the OpenAI Sora API ended on 2026-09-24, according to the Magic Hour release tracker (read 2026-10-06), so many teams are re-picking a model rather than swapping one endpoint. The same tracker lists Seedance 2.5 (2026-07-31) and Kling VIDEO 3.0 (2026-02-05) as current options. Sume's catalog ships both as seedance-2.5 and kling-3; the full id list is in Sume video models.

Why must the split be derived from the job key?

A random coin flip on each submit is the common mistake. It breaks the moment a worker retries after a timeout, because the second attempt can draw the other arm. Derive the arm from something that never changes for the logical job, such as the order id or the shot id.

Hash it with SHA-256, take four bytes, and reduce modulo 100. Python's built-in hash() is salted per process, so it is the wrong tool here; a restart would reshuffle every job. Keep the shares in one list so moving from 50/50 to 90/10 is a one-line change.

What does the code look like?

Sume accepts an Idempotency-Key header on submits. A retry with the same key and the same payload returns the original job, and the same key with a different payload returns 409 idempotency_conflict, per Sume generation admission. Append the model id to the key: if you later change the shares, an old key keeps its old arm instead of conflicting.

import hashlib
ARMS = [("seedance-2.5", 50), ("kling-3", 50)]

def bucket(job_key: str) -> int:
    digest = hashlib.sha256(job_key.encode()).digest()
    return int.from_bytes(digest[:4], "big") % 100

def pick_model(job_key: str) -> str:
    n, edge = bucket(job_key), 0
    for model, share in ARMS:
        edge += share
        if n < edge:
            return model
    return ARMS[-1][0]

def build(job_key: str, prompt: str) -> dict:
    model = pick_model(job_key)
    return {"headers": {"Idempotency-Key": f"{job_key}:{model}"},
            "body": {"model": model, "prompt": prompt, "duration": 8}}

if __name__ == "__main__":
    first = build("order-1001", "A kettle boiling, macro shot")
    assert build("order-1001", "A kettle boiling, macro shot") == first
    counts = {m: 0 for m, _ in ARMS}
    for i in range(1000):
        counts[pick_model(f"order-{i}")] += 1
    print(first["body"]["model"], counts)
    assert all(380 < c < 620 for c in counts.values())

Which parameters can both arms share?

Both arms must be legal for the same request, or one arm fails on validation and your result is about error rates, not quality. Pick a duration inside both ranges. GET /v1/videos/models returns the live catalog, and a request the catalog rejects comes back as 400 unsupported_capability, so check the pair once at deploy time.

Duration limits from Sume video models, for choosing a shared request (read 2026-10-06)
Model idDuration rangeNote
seedance-2.54 to 30 sTracker lists up to 30 s
kling-3up to 15 sTracker lists 15 s with native audio
wan-3.02 to 30 sWidest range of the five
minimax-h35 to 15 s480p or 768p, not 720p
gemini-omni-flash-1.13 to 10 sDefault 8 s

What should stay identical across arms?

Keep everything except model identical between arms: the prompt, the duration, the aspect ratio, and the input images. Sume rejects seed on /v1/videos with 400 unsupported_parameter, and no catalog model accepts it, so you cannot pin the randomness to compare two models on a matched pair. Compare distributions over many jobs instead of pairs of clips, and let a reviewer score each clip without seeing which arm made it.

Do not use sume/auto as an arm. It is a router, not a model: Sume documents its resolution as a pure function of the normalized request and the catalog version, never discloses which family served a request, and echoes sume/auto as the job's model. An arm you cannot name is not a comparison, so pin explicit catalog ids on both sides.

What should you record?

Log the model id, the job id, and your quality score together, and compare the arms only after each has a few hundred completed jobs. Sume does not publish a quality benchmark, and this post does not invent one: your own rubric is the measurement. Cost differs by model, so record the billed amount per job from your usage data rather than from list prices.

Ramp carefully. Start at 10 percent on the new arm, because a new model id can fail validation in ways the old one did not, and a bad arm at 50 percent doubles the blast radius. When the shares change, existing keys keep their arm, since the key embeds the model, so jobs already in flight are unaffected.

If one arm starts returning 429 queue_full, that is capacity, not quality. Drop it from the comparison for that window instead of counting the failures against the model.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume