Sora Batch API render queue gone: Sume async jobs instead

OpenAI's Sora guide listed Batch API support before the September 24 shutdown. On Sume the pattern is async jobs, a queue and signed webhooks.

5 min readSume
All posts

If you queued overnight Sora renders through OpenAI's Batch API, that route ended with the Videos API on September 24, 2026. On Sume the equivalent is not a batch endpoint but a queue: submit many video jobs with mode: "async" or a callback_url, let them wait in queued, and verify the signed webhook when each finishes.

What did OpenAI say about Sora and batches?

OpenAI's video generation guide, read on 2026-10-02, says the Sora 2 models and the Videos API were permanently shut down on September 24, 2026, and that no one-to-one replacement API is available. It lists batch processing among the features that are now unavailable. The deprecations page confirms the shutdown date and lists sora-2, sora-2-pro and their dated snapshots with no replacement.

So there is nothing to migrate to at OpenAI for video. The question is how to rebuild the offline-queue pattern somewhere else.

How does Sume queue work?

Per the generation admission docs, a valid submit creates a durable job. If your workspace is at its processing concurrency, the job is accepted as queued while queue capacity remains, and workers move it to processing later. Concurrency is a dispatch limit, not a submit limit.

Sume plan concurrency and queue capacity (docs read 2026-10-02)
PlanProcessing concurrencyQueue capacityAccepted jobs
Free156
Pro42024
Startup84048
Scale20100120

What happens when the queue is full?

New paid generation submissions fail with 429 queue_full, and request-volume limits fail with 429 rate_limited. Retry with backoff and an idempotency key. A shortfall in balance fails with 402 insufficient_credits before any provider work starts. These are four different controls, and the docs say not to confuse them. Check generation_limits.concurrency_limit for your effective number rather than trusting a static table.

How do I get results without polling?

Pass a public HTTPS callback_url on the video request; Sume signs the raw body with HMAC SHA-256 over <timestamp>.<raw_body> and sends x-sume-webhook-timestamp and x-sume-webhook-signature. The signature is sume-v1=<hex>. The sketch below verifies it, rejects an empty secret and a stale timestamp, and accepts any entry during a secret rotation.

import asyncio, hashlib, hmac, time

def verify(raw: bytes, ts: str, header: str, secret: str) -> bool:
    if not secret or abs(time.time() - int(ts)) > 300:
        return False
    mac = hmac.new(secret.encode(), f"{ts}.".encode() + raw, hashlib.sha256)
    want = "sume-v1=" + mac.hexdigest()
    return any(hmac.compare_digest(e.strip(), want) for e in header.split(","))

async def main():
    secret, raw, ts = "whsec_demo", b'{"event":"job.completed"}', str(int(time.time()))
    sig = "sume-v1=" + hmac.new(secret.encode(), f"{ts}.".encode() + raw, hashlib.sha256).hexdigest()
    print(verify(raw, ts, sig, secret), verify(raw, ts, sig, ""))

asyncio.run(main())

What does Sume not give you?

There is no single batch file upload and no half-price batch tier in the documents I read, and Sume bills list times 1.25 on Video Router models, so do not expect a batch discount. Delivery is up to 10 attempts at a fixed 30-second spacing by default, so store the event and answer 2xx fast, keeping status_url polling as the fallback.

Use job_id as your idempotency key, and read the video docs for each model's supported durations and resolutions.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume