Canary 5-10% of image jobs to a new Sume image model in Python

Route a stable slice of image jobs to a new model by hashing the job key, log the model and cost, and widen the slice only if review scores hold.

5 min readSume
All posts

To canary a new image model, hash a stable job key (an order id, not a random number) into a bucket from 0 to 99, send buckets below your percentage to the new model id, and record the model and the returned usage.cost for every job. Hashing keeps a given job on the same model across retries, so the comparison stays clean and a re-run does not flip between models.

On Sume the change is one string: model in the POST /v1/images body. Everything else in the request can stay the same, provided the new model accepts the parameters you send; a model that does not list one returns 400 unsupported_parameter.

What should the canary compare?

Decide the measures before you route traffic.

Canary measures, read 2026-10-04
MeasureSourceWhy
Billed costusage.cost in the 200 bodySume reports the billed USD amount per call
Error rateHTTP status and the error codeA new model may reject a parameter the old one accepted
LatencyYour own timerA 202 means the job outlived the sync wait
Approval rateYour reviewersCost per approved image is the number that matters

What is the router?

It uses SHA-256 of the job key so the bucket is the same on every machine and every run. Set both model ids to ones you have checked in GET /v1/images/models.

import hashlib, os, requests

STABLE = "google/nano-banana-2"
CANDIDATE = "bytedance-seed/seedream-5-lite"
PERCENT = 10

def bucket(key):
    return int(hashlib.sha256(key.encode()).hexdigest(), 16) % 100

def generate(job_key, prompt):
    model = CANDIDATE if bucket(job_key) < PERCENT else STABLE
    r = requests.post(
        "https://api.sume.com/v1/images",
        headers={"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"},
        json={"model": model, "prompt": prompt},
        timeout=90,
    )
    body = r.json() if r.ok else {}
    cost = body.get("usage", {}).get("cost")
    print(job_key, model, r.status_code, cost)
    return model, r.status_code, body

if __name__ == "__main__":
    generate("order-1042", "A ceramic mug on a wooden table, soft daylight")

When do I widen the slice?

Wait until you have enough approved and rejected images in each arm to see a difference, then step from 10 to 25 to 50 percent. Keep the stable model id in config so rollback is a one-line change. The video version of this pattern is in canarying video jobs.

How do retries interact with the canary?

Retry on the same model the job was bucketed to. Failed or cancelled generations are not billed on Sume, while a completed one is billed in full, so a retry after a timeout can bill twice if the first call actually completed. Read the retry post before wiring automatic retries.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume