OpenAI Tier 1 is 200 requests a minute: poll Sume jobs in batches

A job-polling loop burns a 200 requests-a-minute limit fast. One batched jobs_wait on up to 20 Sume job ids replaces dozens of status calls.

5 min readSume
All posts

Replace per-job polling with one batched jobs_wait. OpenAI's remote MCP guide lists rate limits under its limitations, with Tier 1 at 200 requests per minute and higher tiers up to 2,000 (read 2026-10-05). The guide does not say which request types count, so check that against your own account. Either way, a loop that checks twenty jobs every few seconds spends requests quickly, and Sume's jobs_wait takes 1 to 20 ids in one call.

The arithmetic

Say an agent has 20 video jobs in flight and checks each one every 5 seconds. That is 12 checks a minute per job, so 20 x 12 = 240 calls a minute, which is above 200. Even at one check every 10 seconds the loop makes 120 calls a minute, more than half of the budget, before the model does anything else.

With a batched wait, the agent makes one call per slice. The remote MCP slice defaults to 50 seconds and is capped at 55, so one call covers up to 55 seconds. That is about 1.1 calls a minute at the cap.

Polling load for 20 jobs against a 200 requests-a-minute limit, read 2026-10-05
PatternCalls per minuteShare of 200
Status every 5 s per job240120%
Status every 10 s per job12060%
Status every 30 s per job4020%
One batched jobs_wait, 55 s slicesAbout 1.1About 0.5%

How batched waiting works

The Sume jobs docs say jobs_wait accepts job_id for one job, or job_ids with 1 to 20 ids and wait_for of all or any. The response is a job_wait_batch with a status snapshot for each id. If a slice expires with wait_slice_expired, call again on the same ids and never submit the paid create again.

Pass include_results: true and finished jobs return their jobs_result inline, so a wave needs no separate reads. Results that do not fit come back in results_omitted.job_ids.

Compute your own budget

The helper below turns a poll interval and a job count into calls per minute. It runs offline.

def calls_per_minute(jobs: int, every_s: float) -> float:
    return jobs * 60 / every_s

def batched_calls_per_minute(slice_s: float = 55) -> float:
    return 60 / slice_s

if __name__ == "__main__":
    for every in (5, 10, 30):
        print(every, calls_per_minute(20, every))
    print(round(batched_calls_per_minute(), 2))

Where the requests go

A polling loop is only one source of requests. Every model turn that calls a tool, every retry, and every discovery call such as tools_list adds to the same minute. Count the whole agent, not only the wait loop, when you compare against a limit.

The practical rule is to let one call carry as much as it can. Submit a wave sized to your plan's headroom, wait on all ids at once, and read results inline with include_results. That shape gives the agent few, large calls, which suits a tight request limit and a slow render equally well.

Checklist

Make polling cheap before it is a problem.

  • Use jobs_wait with job_ids, not one status call per job.
  • Treat a 524 on jobs_wait as transport failure, then wait again on the same ids.
  • Cap the number of jobs in flight using generation_limits headroom.
  • Never resubmit a paid job because a wait slice ran out.

Sources

Related posts

More in Agents

All Agents posts

Written by Sume