Format batch of 100 runs: writes bind first, reads do not (by plan)

Format runs poll cheaply: about one read a minute per run in flight. Creates are the scarce budget, 120 a minute on Free. Poll math for a 100-run batch by plan.

6 min readSume
All posts

If you poll Format runs with the documented backoff, each run in flight costs about one read a minute once the gap reaches 60 seconds, so 100 runs in flight is roughly 100 reads a minute against a read budget of 4,800 on Free. The scarce budget is writes, because creating a run is a write: 120 a minute on Free, 300 on Pro, 600 on Startup and 1,200 on Scale.

The budgets come from Format errors and spend, the polling loop from Format runs and results, and the 15 to 30 minute typical time for long-form video from the Format API intro. The per-run read count below is arithmetic on the documented loop, not a measured figure.

What are the budgets by plan?

Every key has a per-minute request budget across all of /v1, set by the workspace's plan. Reads and writes are separate, and the read budget is forty times the write one. A read is any GET: the receipt, status_url, events_url, result_url, and the Format and run lists. A write is everything else, including creating runs and queues, cancel and redeliver.

Per-minute request budgets by plan (read 2026-10-03)
PlanWrites per minuteReads per minuteProcessing concurrency
Free12048001
Pro300120004
Startup600240008
Scale12004800020

How many reads does one run cost?

The documented loop sleeps 5 seconds, then doubles the gap up to 60. For a run that takes 25 minutes, the first four reads land at roughly 5, 15, 35 and 75 seconds, and after that the loop reads once a minute. That is about 4 reads in the first minute and a quarter, then about 24 more over the remaining 23 minutes and 45 seconds: roughly 28 reads for the whole run, or one a minute in steady state.

At that rate, 100 runs in flight is about 100 reads a minute in steady state, with a burst while they all start. Even the Free read budget of 4,800 a minute holds that with room to spare, and the docs say a poll loop cannot starve your own creates because the budgets are separate.

  • Poll status_url for the small payload while waiting; it carries status, next_action, expires_at and queue.
  • Read result_url or the receipt once, after the status is terminal.
  • Use expires_at as your ceiling: a non-terminal run is force-finalized as failed 90 minutes after creation, or sooner if it has gone silent.

Where the batch actually waits

Creating 100 runs is 100 writes, which a Free key can do inside one minute, and a bulk queue accepts up to 100 items in a single request. The wait is elsewhere: how many generations run at once is governed by the plan's concurrency limit, and raising your request rate does not raise it. On Free that limit is 1, on Pro 4, so a large share of a 100-run batch can sit queued or processing while it waits for a slot (a bulk queue also takes its own required concurrency window of 1 to 16), and the right answer is to submit once and poll slowly.

A 429 during polling is transient. The run is still executing and still spending, so wait for retry-after and read again instead of treating it as a failure. The response headers ratelimit-limit, ratelimit-remaining and ratelimit-reset tell you how much budget is left, and error.details.scope says whether the read or write budget ran out.

A polling helper that follows the headers

This Python helper reads status_url, honors retry-after on a 429, doubles the gap up to 60 seconds, and gives up at your own ceiling. Set the ceiling from expires_at when you have it.

import os
import time
import requests

BASE = "https://api.sume.com"
HEAD = {"Authorization": "Bearer " + os.environ["SUME_API_KEY"]}


def wait_for_run(run_id, ceiling_seconds=5400):
    gap = 5
    started = time.time()
    while time.time() - started < ceiling_seconds:
        r = requests.get(
            f"{BASE}/v1/format-runs/{run_id}/status", headers=HEAD, timeout=30
        )
        if r.status_code == 429:
            time.sleep(int(r.headers.get("retry-after", "5")))
            continue
        r.raise_for_status()
        data = r.json()["data"]
        if data["status"] not in ("queued", "processing"):
            return data
        time.sleep(gap)
        gap = min(gap * 2, 60)
    raise TimeoutError(f"run {run_id} still in flight at your ceiling")

Sources

Related posts

More in Formats

All Formats posts

Written by Sume