Sume video mode: async, sync, subscribe or webhook? Cost is the same

Mode only decides how you learn the outcome: all four create the same job at the same price. A decision table for web apps, workers and batch pipelines.

4 min readSume
All posts

Pick async for almost everything, add a callback_url when you can receive HTTPS, and treat sync and subscribe as a short convenience wait. The Sume docs are explicit that mode never changes whether a job is created, what it costs, or how long it runs; it only decides how you learn the outcome.

If you omit mode you get async. If you send webhook_url, or its alias callback_url, without a mode, you get webhook.

The four modes

From the jobs and results page. "Server blocks" is the part that surprises people: it is at most 30 seconds and can be less.

Communication modes for generation submits (read 2026-10-07)
ModeHTTP returnsServer blocksClient does next
async202 with job envelopeNoPoll status_url until terminal, then read result_url
syncEnvelope after up to 30 sYes, at most 30 sTerminal: read it. Not terminal: poll, do not resubmit
subscribeSame as syncSame as syncSame as sync
webhook202 with job envelopeNoWait for the signed callback; keep polling as backup

Choosing by app shape

A web app with a job table and a worker should use async or webhook and never hold a browser request open on a video. A serverless function with a hard time limit should do the same, because a video usually runs longer than any wait Sume offers.

sync is useful for image calls and short tests, where the result often fits inside the budget. When it does not, the response has sync.timed_out: true, and the right move is to poll the job that already exists. sync.capacity_exhausted: true means Sume skipped the wait because the process's waiter budget was full.

  • Always store the job id from the first response.
  • Webhook mode delivers only terminal events: job.completed, job.failed, job.canceled. There are no progress or partial webhooks.
  • Keep status_url polling as the backup in every mode.
  • There is no SSE stream for generation jobs; for a fal-style long wait, use async and the SDK's waitForJob.

What a retry looks like

A client-side timeout never cancels the job. Retrying the submit with the same Idempotency-Key returns the original job, so the mode you chose does not add risk of duplicates. A new key, in any mode, is a new job and a new charge.

curl -s https://api.sume.com/v1/videos \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Idempotency-Key: clip-8823-v1" \
  -H "Content-Type: application/json" \
  -d '{"model":"sume/auto","prompt":"Slow push-in on a ceramic mug","callback_url":"https://example.com/hooks/sume"}'

A sensible default

Start with async plus a signed callback_url, and a sweeper that polls any job older than your expected render time. That gives you push when the network is healthy and a poll when it is not, which is what the docs recommend for production.

Switch to sync only for a demo or a test script where blocking up to 30 seconds is acceptable. Never rely on it for video: a job that outlives the wait returns a non-terminal envelope, and the right response is to poll the existing job.

  • Store the job id before anything else.
  • Verify the callback signature on the raw body.
  • Use one Idempotency-Key per logical request in every mode.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume