Canary 5-10% of image jobs to a new Sume image model in Python
Route a stable slice of image jobs to a new model by hashing the job key, log the model and cost, and widen the slice only if review scores hold.

To canary a new image model, hash a stable job key (an order id, not a random number) into a bucket from 0 to 99, send buckets below your percentage to the new model id, and record the model and the returned usage.cost for every job. Hashing keeps a given job on the same model across retries, so the comparison stays clean and a re-run does not flip between models.
On Sume the change is one string: model in the POST /v1/images body. Everything else in the request can stay the same, provided the new model accepts the parameters you send; a model that does not list one returns 400 unsupported_parameter.
What should the canary compare?
Decide the measures before you route traffic.
| Measure | Source | Why |
|---|---|---|
| Billed cost | usage.cost in the 200 body | Sume reports the billed USD amount per call |
| Error rate | HTTP status and the error code | A new model may reject a parameter the old one accepted |
| Latency | Your own timer | A 202 means the job outlived the sync wait |
| Approval rate | Your reviewers | Cost per approved image is the number that matters |
What is the router?
It uses SHA-256 of the job key so the bucket is the same on every machine and every run. Set both model ids to ones you have checked in GET /v1/images/models.
import hashlib, os, requests
STABLE = "google/nano-banana-2"
CANDIDATE = "bytedance-seed/seedream-5-lite"
PERCENT = 10
def bucket(key):
return int(hashlib.sha256(key.encode()).hexdigest(), 16) % 100
def generate(job_key, prompt):
model = CANDIDATE if bucket(job_key) < PERCENT else STABLE
r = requests.post(
"https://api.sume.com/v1/images",
headers={"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"},
json={"model": model, "prompt": prompt},
timeout=90,
)
body = r.json() if r.ok else {}
cost = body.get("usage", {}).get("cost")
print(job_key, model, r.status_code, cost)
return model, r.status_code, body
if __name__ == "__main__":
generate("order-1042", "A ceramic mug on a wooden table, soft daylight")When do I widen the slice?
Wait until you have enough approved and rejected images in each arm to see a difference, then step from 10 to 25 to 50 percent. Keep the stable model id in config so rollback is a one-line change. The video version of this pattern is in canarying video jobs.
How do retries interact with the canary?
Retry on the same model the job was bucketed to. Failed or cancelled generations are not billed on Sume, while a completed one is billed in full, so a retry after a timeout can bill twice if the first call actually completed. Read the retry post before wiring automatic retries.
Sources
Related posts
More in Developers
- Cancel a Sume Omni job: only before it starts, then 409
POST /v1/jobs/{id}/cancel works while a Sume job is queued. Once generation starts you get 409 job_generation_already_started and the clip is still billed.
- Can you cancel a Sume TTS job? Only before generation starts
Sume cancels a job only before generation work begins. After that the cancel call is rejected with 409. How to code for both outcomes.
- Captions for a clip that switches English and Spanish mid-sentence
MAI-Transcribe-2-Streaming advertises continuous language detection. For a recorded code-switching clip on Sume, lock the wording with script_text.
- Watch the Sume video catalog for new ids and changed limits in Python
Fetch GET /v1/videos/models, save a snapshot and diff new ids, removed ids and changed durations or resolutions. A Python script, testable offline.
Written by Sume