GPT-6.1 Sol rate limits (Tier 1: 500 RPM) vs a Sume bulk run window

OpenAI lists GPT-6.1 Sol limits from 500 RPM at Tier 1 to 15,000 RPM at Tier 5. A Sume bulk run uses a concurrency window of 1-16 and 100 items.

4 min readSume
All posts

OpenAI's model page lists GPT-6.1 Sol rate limits by usage tier, from 500 requests per minute and 500,000 tokens per minute at Tier 1 to 15,000 requests per minute and 40,000,000 tokens per minute at Tier 5. Those limits belong to an OpenAI account calling the OpenAI API. A Sume bulk run has its own controls: a concurrency window of 1 to 16 and between 1 and 100 items.

The two do not describe the same thing, and this post does not claim how one maps to the other. It sets out each so you can plan around the one you actually call.

What are the vendor limits?

The model page's rate-limit table gives requests per minute (RPM) and tokens per minute (TPM) for each tier.

GPT-6.1 Sol rate limits, OpenAI model page, read 2026-10-03
TierRPMTPM
Tier 1500500,000
Tier 25,0001,000,000
Tier 35,0002,000,000
Tier 410,0004,000,000
Tier 515,00040,000,000

What does a Sume bulk run limit?

The bulk runs doc defines a queue of up to 100 Format runs. concurrency is a required integer from 1 to 16 and says how many child runs stay in flight at once; items is a required array of 1 to 100 entries, each the same body as a single run. Because each item is a full run body, each can carry its own model.

The doc says workspace generation concurrency still applies on top of the window, and that a 429 rate_limited means wait retry-after seconds. Creating the queue spends the write budget; polling spends a separate, larger read budget, so a poll loop does not starve your creates.

Do the vendor tiers apply to my Sume run?

The Sume docs do not say. The vendor table is the limit of an OpenAI account; the Sume docs describe the limits of the Sume API. Read the 429 behavior in the call doc and plan to Sume's retry-after, not to the vendor's RPM. If you also call OpenAI directly in the same pipeline, the RPM and TPM columns above are the budget for that part.

A small sizing example

Say you plan 60 Format runs and each takes about two minutes. With concurrency 6, ten rounds of six runs would take roughly twenty minutes; with 16 it is under four rounds, so about eight minutes. The numbers are an illustration of the window, not a measured run time. The real time depends on the Formats and on workspace generation concurrency.

Raise the window only as far as your workspace allows and back off on a 429.

What does the vendor page add beyond rate limits?

The same OpenAI page lists a 1,050,000-token context window, a 128,000-token maximum output, and says tool calling requires the Responses API. It also says the prices differ between prompts under and over 272K input tokens. For a pipeline that sends many large prompts, the TPM column often binds before the RPM column does: at Tier 1, a 500,000 TPM budget is exhausted by two requests of 250,000 tokens in a minute, long before 500 requests are made.

That arithmetic applies to a direct OpenAI integration. For Sume, the unit to plan around is the Format run and its receipt.

How should a pipeline handle both?

Treat the two limit systems as independent until a doc says otherwise.

  • Use the Sume concurrency window to bound Format runs in flight; start low and raise it.
  • Honor retry-after on any 429 from the Sume API, and use a separate back-off for any direct OpenAI calls.
  • Poll the queue at a modest interval; polling has its own read budget.
  • Record the echoed model and usage fields from each receipt so cost and model changes are visible.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume