Sora to Sume in Python: a 10% rollout flag with a spend guard
Move video traffic off a dead Sora call one slice at a time. A stable per-user percentage flag, one Sume call, and a cost guard that stops at your daily cap.

OpenAI's Sora API is reported shut down (Pinggy, read 2026-10-07), so any code path that still calls it is failing. If your users are on a feature flag already, the safest move is not a big-bang swap but a percentage rollout to Sume, with a spend cap that stops the rollout before a bug becomes a bill.
The function below hashes the user id to a stable bucket, so one user always lands on the same side. It submits to POST /v1/videos with model: "sume/auto", polls the job, and returns the download URL together with usage.cost.
The code
Set SUME_API_KEY and SUME_PERCENT. Python 3.9 or newer with requests installed.
import os, time, zlib, requests
BASE = "https://api.sume.com/v1/videos"
HDR = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
PERCENT = int(os.environ.get("SUME_PERCENT", "10"))
DAILY_CAP_USD = 25.0
spent = 0.0
def on_sume(user_id: str) -> bool:
return zlib.crc32(user_id.encode()) % 100 < PERCENT
def render(prompt: str, seconds: int = 8) -> tuple[str, float]:
global spent
if spent >= DAILY_CAP_USD:
raise RuntimeError("daily cap reached")
body = {"model": "sume/auto", "prompt": prompt, "duration": seconds, "aspect_ratio": "9:16"}
job = requests.post(BASE, headers=HDR, json=body, timeout=60)
job.raise_for_status()
url = job.json()["polling_url"]
while True:
s = requests.get(url, headers=HDR, timeout=60).json()
if s["status"] == "completed":
cost = float(s["usage"]["cost"])
spent += cost
return s["unsigned_urls"][0], cost
if s["status"] in ("failed", "cancelled"):
raise RuntimeError(s.get("error", "job failed"))
time.sleep(30)Why these choices
sume/auto accepts 3 to 10 seconds at 16:9 or 9:16 and resolves to the same model every time, so a replay prices identically. Its defaults are 720p and 8 seconds, which is $1.00 per clip at the Omni Flash rate, so a $25 daily cap allows 25 clips. Raise SUME_PERCENT to 25, 50 and 100 only after you have compared the first batch with the old output.
The guard is per process. If you run several workers, keep the running total in a shared store. And cap by reserve, not just by finished cost: a job in flight has already reserved its price.
The rollout order
Start with an internal account at 100 percent, then 10 percent of users for a day, then 50, then everyone. Keep the Sora branch only as a deleted code path; there is no old API left to fall back to, so the fallback is a different Sume model id, not a rollback.
Sources
Related posts
More in Developers
- Speaking rate in words per minute from Sume STT word times (Python)
Compute words per minute for a recording from the words[] start and end times Sume STT returns, plus a per-minute pacing table. Offline Python, no API call.
- Speech to text API in Go: transcribe audio with net/http
Transcribe audio in Go using only the standard library: submit to Sume STT, poll the job and print the text. A 30-line program at one cent per audio minute.
- Speech to text API in Node.js: transcribe audio with fetch
Transcribe audio in Node.js with built-in fetch: submit to Sume STT, poll the job and print sentence segments with timestamps. 30 lines, no dependencies.
- Speech to text API in Ruby: transcribe audio with Net::HTTP
Transcribe audio in Ruby with only the standard library: submit to Sume STT, poll the job, print text and word times. A 30-line script at one cent a minute.
Written by Sume