Luma Agents API polling: 20 s wait, 10-minute video timeout, and Sume
Luma's quickstart says wait 20 seconds, then poll, with a hard 2-minute image and 10-minute video timeout, not backoff. A matching Sume loop that actually runs.

Luma's Agents API quickstart recommends waiting about 20 seconds before the first poll, because the uni-1 median is around 30 seconds, and setting a hard timeout of 2 minutes for images and 10 minutes for video rather than using exponential backoff. Sume's own video example polls every 30 seconds, and its jobs docs tell clients to follow next_poll_after_seconds when it is present.
Luma's page is its quickstart, read 2026-10-03. Sume's are the video generation docs and jobs and results.
What exactly does Luma recommend?
The workflow is three steps: submit to POST /v1/generations, poll GET /v1/generations/{generation_id}, and download from a presigned URL once the state is completed. For production, the quickstart suggests a 20-second initial wait, then a fixed hard timeout instead of backing off. It gives 2 minutes for images and 10 minutes for video.
The reasoning is stated for images: the uni-1 p50 is about 30 seconds, so polling earlier is wasted calls. The guide does not publish a p50 for ray-3.2, so the 10-minute video figure is a ceiling to give up at, not a typical wait.
| Item | Luma Agents API | Sume |
|---|---|---|
| First wait | About 20 seconds | The video example sleeps 30 seconds before each poll |
| Interval | Not specified beyond the first wait | next_poll_after_seconds when present, else backoff (jobs docs) |
| Give up | 2 minutes images, 10 minutes video | Not fixed; your client chooses the timeout |
| Push option | Poll-based per the migration guide | callback_url, HTTPS, signed with x-sume-webhook-signature |
What do Sume's docs say about waiting?
Two things that look like they disagree but do not. The video page's sample loop sleeps a flat 30 seconds between polls. The jobs page says to use exponential backoff and, when the status carries next_poll_after_seconds, to follow that value instead. A sync-mode wait is capped at 30 seconds, and a job that is not terminal then has to be polled, not resubmitted.
The jobs page also says the wait for anything that can outlast 30 seconds lives in your client, so its timeout can be as long as you need. That is the part Luma's quickstart supplies as a number: pick one, and 10 minutes is a defensible starting point for a clip.
What does a matching Sume loop look like?
This submits a short clip, polls the returned URL every 30 seconds, and stops at 10 minutes, mirroring Luma's ceiling. It uses the statuses from the video docs.
import os, time, requests
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
r = requests.post("https://api.sume.com/v1/videos", headers=H, json={
"model": "seedance-2.5", "prompt": "A slow dolly past a rain-wet street",
"resolution": "480p", "duration": 5})
r.raise_for_status()
url = r.json()["polling_url"]
deadline = time.time() + 600
while time.time() < deadline:
time.sleep(30)
s = requests.get(url, headers=H).json()
if s["status"] == "completed":
print(s["unsigned_urls"][0], s["usage"]["cost"])
break
if s["status"] in ("failed", "cancelled"):
print("stopped:", s.get("error"))
break
else:
print("gave up; job still running:", url)What happens at the timeout?
Giving up on the client does not cancel the job on either side as far as these pages say. The loop above prints the polling URL so you can resume later. On Sume, the same job is also visible at GET /v1/jobs/{id}/status, so a timed-out client can pick it up again without resubmitting, and sending an Idempotency-Key makes a retried submit return the original job instead of a second one.
Why not back off exponentially?
Backoff helps when load is unknown and a slow poll costs little. Luma's guidance points the other way for generation: the first answer is rarely ready before about 30 seconds, so a long initial wait followed by steady polling and a firm timeout is simpler and bounds your worst case. Sume's jobs page does recommend backoff, and also hands you next_poll_after_seconds so the server can set the pace. Both are reasonable; the thing to avoid is a tight loop with no timeout at all, which on either API is how a stuck client turns into a rate-limit problem.
Sources
Related posts
More in Developers
- Luma Ray 2 to Ray 3.2 migration: the field map and what changes
Luma is retiring Ray 3, Ray 2, Ray 2 Flash and more for ray-3.2 on one /v1/generations endpoint. The field map, the poll-only change, and Sume's contrast.
- MAI-Transcribe-2-Streaming Realtime API: events vs Sume job URLs
Microsoft's streaming transcriber uses a WebSocket with delta, intermediate and commit events. Sume STT takes a file URL and returns a job. A side-by-side.
- Migrate Video 1.0 to sume/auto on POST /v1/videos, field by field
Video 1.0 is retiring soon. Which fields move, which are ignored or rejected, and the request to send on POST /v1/videos with sume/auto.
- Mix Seedance, Kling and Omni clips in one video: shared aspect ratio
16:9 and 9:16 are the aspect ratios Seedance 2.5, Kling 3 and Gemini Omni Flash 1.1 all list. Set the Timeline output to match and plan before render.
Written by Sume