Streamlit Sora demo: switch to Sume and st.video, no double billing
Streamlit reruns your script on every click. Derive the Sume Idempotency-Key from the prompt so a rerun returns the first job, then show the clip with st.video.

A Streamlit app that called Sora reruns its whole script on every widget change, so a naive port to Sume can submit the same billable video twice. Derive the Idempotency-Key from the prompt, so a rerun with the same text gets the first job back, then fetch the bytes from /v1/videos/{id}/content and pass them to st.video, because the content URL needs your API key and a browser cannot send it.
OpenAI's deprecations page, read 2026-10-08, lists the Videos API as removed on September 24, 2026, so a demo that still calls it is already broken.
Why a prompt-derived key
Sume answers 409 idempotency_conflict when a key is reused with a different body. That works in your favour: the same prompt and the same fields hash to the same key and replay the first job; a changed prompt hashes to a new key and starts a new job. If a user wants a second take of the same text, add a take number to the hashed string.
| User action | Script reruns | Key result | Outcome |
|---|---|---|---|
| Click Generate with new text | yes | new hash | new job, reserve list price x 1.25 |
| Change an unrelated widget | yes | none, the button is not pressed | no request is sent |
| Click Generate again with the same text | yes | same hash | first job returned |
| Same key, different body | yes | same key | 409 idempotency_conflict |
The app
Run it with streamlit run app.py and SUME_API_KEY set. The call blocks the script thread while it polls, which is acceptable for a demo; use a worker and a webhook for anything with real traffic.
import hashlib, os, time
import requests
import streamlit as st
BASE = "https://api.sume.com"
HEAD = {"Authorization": "Bearer " + os.environ["SUME_API_KEY"]}
def generate(prompt):
key = hashlib.sha256(prompt.encode()).hexdigest()[:32]
res = requests.post(BASE + "/v1/videos", headers={**HEAD, "Idempotency-Key": key},
json={"model": "sume/auto", "prompt": prompt}, timeout=60)
res.raise_for_status()
job = res.json()
while job["status"] in ("pending", "in_progress"):
time.sleep(30)
job = requests.get(BASE + "/v1/videos/" + job["id"], headers=HEAD, timeout=60).json()
if job["status"] != "completed":
return None
url = BASE + "/v1/videos/" + job["id"] + "/content?index=0"
return requests.get(url, headers=HEAD, timeout=300).content
prompt = st.text_input("Prompt")
if st.button("Generate") and prompt:
with st.spinner("Rendering, this can take a minute or more"):
data = generate(prompt)
if data:
st.video(data)
else:
st.error("The job did not complete")Keep the key off the page
Never pass the content URL to st.video or an HTML tag in the browser; the request needs your Bearer token, and the token must stay server side. Save the bytes to disk or object storage if the user may come back later, as the storage post recommends.
Sources
Related posts
More in Developers
- sume/auto and idempotent retries: same price, same route on replay
New video models land weekly. With model sume/auto and an Idempotency-Key, a retry of a Sume video submit gets the original job, price and route.
- model sume/auto on /v1/videos vs pinning Seedance 2.5: what differs
sume/auto lets Sume pick the video family and never says which. Its documented limits are 3 to 10 s, so a 15 or 30 second Seedance clip must be pinned.
- Cancel a Sume bulk run: no queue endpoint, so cancel each child
Sume has no public cancel for a bulk queue. Poll the queue, then POST /v1/format-runs/{run_id}/cancel for each running child. fal cancels one request with PUT.
- Resume a bulk run after a crash: replay the Idempotency-Key
If your client dies after POSTing a 100-item bulk run, replay the same key and same body: Sume returns 202 with the original queue. A new body gets 409.
Written by Sume