Will polling video jobs trigger a 429? Reads and writes differ

Each Sume API key has a write budget and a read budget 40 times larger, so a status-poll loop cannot 429 your submits. Poll math, headers and a 429 backoff.

5 min readSume
All posts

Polling will not eat your submit budget. Each Sume API key gets a per-minute budget for all of /v1, with reads and writes counted separately, and the read budget is forty times the write number. A GET of status_url is a read.

A 429 on a read still means you should slow down. Honor retry-after when it is present, and prefer the next_poll_after_seconds the status response gives you.

The budgets

The plan of the workspace that owns the key sets both numbers. Enterprise is not self-serve. Until Sume provisions a contracted number it uses the Scale row.

Request budget per API key per minute (read 2026-10-07)
PlanWritesReads
Free1204800
Pro30012000
Startup60024000
Scale120048000

Poll math for a batch

Twenty jobs polled every 2 seconds is 600 reads a minute, well under the Free read budget of 4800. A 30-second video usually runs for minutes, so a 10-second interval cuts that to 120 a minute with no loss. The real limit on video throughput is generation concurrency, not request rate, and the docs say plainly that HTTP limits are not generation capacity.

Add jitter so several workers that started together do not stay in phase.

import os, random, time, requests
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}

def read_status(job_id: str) -> dict:
    for attempt in range(6):
        r = requests.get(f"https://api.sume.com/v1/jobs/{job_id}/status",
                         headers=H, timeout=30)
        if r.status_code == 429:
            wait = float(r.headers.get("retry-after", 2 ** attempt))
            time.sleep(wait + random.random())
            continue
        r.raise_for_status()
        return r.json()["data"]
    raise RuntimeError("status reads kept returning 429")

d = read_status("job_123")
print(d["sume_status"], d["next_poll_after_seconds"])

Do not confuse the two 429s

rate_limited is request volume. queue_full means the workspace used all its accepted generation capacity, and it clears when a job finishes or you cancel one. The retry advice differs, but both should reuse the same Idempotency-Key on a submit.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume