Canary 10% of video jobs to Sume before cutover: sticky bucketing
Moving video traffic off a shut-down API: hash a stable key into a percent bucket so each customer stays on one backend, and raise Sume's share in steps.

Route a fixed percentage of video jobs to Sume by hashing a stable key, such as the customer id, into a 0-99 bucket and sending buckets below the threshold to the new backend. The same customer always lands on the same side, so you compare like with like and a retry never hops backends mid-job. This matters most for teams whose previous vendor is gone: OpenAI lists the Sora API shutdown date as 2026-09-24 (read 2026-10-03), so the canary compares Sume against a fallback rather than against the old vendor.
Sticky bucketing
Hash the key, not a random number per request. A random draw would send one customer's retry to a different backend and defeat the idempotency key, which only dedupes within one API.
import asyncio
import hashlib
def bucket(key: str) -> int:
return int(hashlib.sha256(key.encode()).hexdigest()[:8], 16) % 100
def backend(customer_id: str, sume_percent: int) -> str:
return "sume" if bucket(customer_id) < sume_percent else "fallback"
async def main() -> None:
for pct in (10, 50, 100):
share = sum(backend(f"cust-{i}", pct) == "sume" for i in range(1000))
print(pct, share)
asyncio.run(main())What to compare during the ramp
Compare terminal outcome rate, time to terminal and cost per accepted clip, and treat a quality review by a person as a metric too. The Sume side already reports status and error codes in a fixed vocabulary, so the comparison needs only that your fallback adapter maps its own errors into the same buckets.
| Step | Traffic to Sume | Hold until |
|---|---|---|
| 1 | 10% | No spike in failed jobs or 402s |
| 2 | 50% | Queue wait fits your deadline |
| 3 | 100% | Fallback adapter kept, not deleted |
Keep the fallback
Reaching 100% should not delete the interface from the previous section. A vendor that shut down once is the best argument for keeping a second adapter warm.
Sources
Related posts
More in Developers
- Cancel a GPT Image 2.5 job: only possible before it starts
Sume's POST /v1/jobs/{id}/cancel works only before generation work starts. How to read cancelable, what the 409 means, and what a client timeout does not do.
- Chinese text to speech API: set language zh or it reads as English
Sume TTS only guesses Korean and Japanese when the language is missing. For Mandarin send language zh and pick a voice tagged zh, then test one line.
- Choose an AI video model in Python: filter the Sume model list by need
Read GET /v1/videos/models and keep only the models that fit your clip length, resolution and reference needs. A runnable Python filter for Sume.
- allowManagedModsOnly in Claude Code: does hosted Sume MCP still load?
allowManagedModsOnly keeps users' own Claude Code mods from loading. What it leaves alone, how a policy mod reviews the rest, and the Sume MCP connection.
Written by Sume