Measure your own render time on Sume from job event timestamps
No published Sora timing carries over. Read GET /v1/jobs/{id}/events, diff created, started and completed, and keep your own queue and render split per model.

To know how long a Sume video takes for your prompts, subtract timestamps from GET /v1/jobs/{id}/events: job.created to job.started is queue wait, and job.started to job.completed is render time. Keep both per model, because a Sora latency figure from your old logs says nothing about a different model on a different provider.
I make no timing claim for either service here. OpenAI's video guide, read 2026-10-08, only tells you to poll every 10 to 20 seconds, and the Sume docs recommend 30 seconds; neither page promises a render duration.
What each gap means
The events endpoint is a pull snapshot, not a stream, so call it once after the job is terminal. Each event has type and created_at in date-time form.
| Gap | From event | To event | Tells you |
|---|---|---|---|
| Queue wait | job.created | job.started | plan concurrency and queue pressure |
| Submit lag | job.started | generation.submitted | Sume's time to hand work to the model |
| Render | generation.submitted | job.completed | model time for this prompt |
| End to end | job.created | job.completed | what your user waited |
The script
It needs Python 3.11 or later, because datetime.fromisoformat there reads the trailing Z of a date-time string. Set JOB_ID and SUME_API_KEY. It uses the first event of each type and skips a gap when an event is missing, which happens for failed jobs.
import json, os, urllib.request
from datetime import datetime
def when(text):
return datetime.fromisoformat(text)
def main():
req = urllib.request.Request(
"https://api.sume.com/v1/jobs/" + os.environ["JOB_ID"] + "/events",
headers={"Authorization": "Bearer " + os.environ["SUME_API_KEY"]})
events = json.load(urllib.request.urlopen(req))["data"]["events"]
first = {}
for e in events:
first.setdefault(e["type"], when(e["created_at"]))
pairs = [("queue", "job.created", "job.started"),
("render", "generation.submitted", "job.completed"),
("total", "job.created", "job.completed")]
for name, a, b in pairs:
if a in first and b in first:
print(name, (first[b] - first[a]).total_seconds(), "s")
main()Using the numbers
GET /v1/jobs/{id}/status also returns the job with created_at, started_at and completed_at, so the queue and total gaps need one call if you do not want the events.
Run twenty representative prompts per candidate model and keep the median and the slowest, not the average. The queue gap depends on your plan: the Sume jobs docs list concurrency, queue and accepted capacity of 1, 5 and 6 on Free, 4, 20 and 24 on Pro, and 8, 40 and 48 on Startup, so a burst on a small plan shows up as queue time and not as slow rendering.
Feed the result into the replacement scorecard, and use the render median to set your poll sleep and your timeout. The events post explains the individual event types.
Sources
Related posts
More in Developers
- Sora to Sume pull request review: eight lines to check before merge
A reviewer's checklist for a Sora-to-Sume PR: duration and resolution types, status words, the 302 download, the webhook body, idempotency and removed fields.
- Sora API gone: port to Sume with Python urllib, no extra packages
Replace the OpenAI videos calls with a stdlib-only Python script: submit to Sume /v1/videos, poll, then fetch the 302 artifact URL without sending your key.
- Speech-to-text for short voice notes: one request, sync mode
Send a voice note's public URL to Sume STT with mode sync and a 30-second wait. A cent per minute, plus the poll fallback when the job outlasts the wait.
- Streamlit Sora demo: switch to Sume and st.video, no double billing
Streamlit reruns your script on every click. Derive the Sume Idempotency-Key from the prompt so a rerun returns the first job, then show the clip with st.video.
Written by Sume