Polling Gemini Omni jobs on Sume: statuses, backoff, Python example

Poll GET /v1/videos/{id} with backoff: pending, in_progress, completed, failed, cancelled. Python asyncio example for gemini-omni-flash-1.1 from 360p to 4K.

5 min readSume
All posts

Poll the polling_url that POST /v1/videos returns, wait a few seconds between polls, back off to about 30 seconds, and stop on completed, failed or cancelled. For Gemini Omni Flash 1.1 the same loop serves every resolution from 360p to 4K; only the time you wait changes.

The Video generation page documents the job flow: submit, receive a job id and polling URL immediately, poll until the status is completed, then download from the content URL. Its own sample waits 30 seconds between polls.

Statuses to handle

The poll response maps Sume job states to OpenRouter-shaped ones. Handle all five. expired is in the enum for wire compatibility, and Sume does not emit it.

Poll statuses on GET /v1/videos/{id} (docs, read 2026-10-05)
Sume job statusPoll `status`What you do
queuedpendingKeep polling
processingin_progressKeep polling
completedcompletedDownload unsigned_urls[0]
failedfailedRead error; do not retry blindly
canceledcancelledStop

A loop that backs off

This example submits a 3-second 360p job, which is the cheapest valid Omni clip at $0.12, and polls with a delay that grows by half each time up to 30 seconds. asyncio.to_thread keeps the blocking requests call off the event loop so the same pattern works when you poll a batch.

import asyncio, os, requests

H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
URL = "https://api.sume.com/v1/videos"

async def main():
    body = {"model": "gemini-omni-flash-1.1", "prompt": "A paper boat on a puddle, slow push-in",
            "resolution": "360p", "duration": 3, "aspect_ratio": "16:9"}
    r = requests.post(URL, headers={**H, "Idempotency-Key": "poll-demo-001"}, json=body)
    r.raise_for_status()
    job, delay = r.json(), 5.0
    while True:
        s = (await asyncio.to_thread(requests.get, job["polling_url"], headers=H)).json()
        print(s["status"], round(delay, 1))
        if s["status"] in ("completed", "failed", "cancelled"):
            break
        await asyncio.sleep(delay)
        delay = min(delay * 1.5, 30.0)
    print(s.get("unsigned_urls") or s.get("error"))

asyncio.run(main())

Pick the first delay by tier

Higher resolutions and longer clips take longer. The docs give no time per tier, so do not hard-code one. Start with a short first delay for 360p drafts, where you want the answer fast, and a longer one for 4K finals, where polling often is wasted traffic. Measure your own p50 per tier for a day and tune from that.

  • 360p draft, 3 to 5 s: first delay 5 s, cap 15 s.
  • 720p and 1080p: first delay 10 s, cap 30 s.
  • 4K: first delay 15 s, cap 30 s.
  • Treat these numbers as a starting point of your own, not a Sume figure.

After completion

Download from unsigned_urls[0]. If you call the content URL while the job is still running, you get 409 job_not_completed, which is retryable; after a terminal failure you get 409 job_failed, which is not. A 429 rate_limited means slow down your polling and submits.

Keep the ids

Log the job id with the resolution, the prompt and the idempotency key at submit time. A poll loop that crashes mid-way can restart from the id; the job keeps running on the server either way. Send an Idempotency-Key with every submit so a retried request does not queue a second job; reusing a key with a different body returns 409.

Polling a batch

To poll many jobs, keep one list of pending jobs and one loop that checks each in turn, rather than one loop per job. Spread the checks so you do not send a burst every few seconds: with 24 pending jobs and a 30 second cap, one request every 1.25 seconds is a steady rate. Remove a job from the list when it hits a terminal status, and write its result to your store at once.

Sume returns 429 rate_limited when you exceed your request budget, so a polling loop that is too eager can crowd out your own submits. Treat 429 as a signal to double the delay.

Cost of polling

The money is spent at submit, not at poll: the balance is reserved when the job is created. The costly mistake is submitting twice. Use idempotency keys on every submit, and log the polling URL so you can resume after a restart without a new submit.

When polling is the wrong tool

A big batch of 4K jobs spends most of its polls on in_progress. A webhook sends one request per job at the end. See Webhooks and Jobs and results for the four communication modes and the poll fallback to keep beside a webhook.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume