GPT-6.1 Sol rate limits (Tier 1: 500 RPM) vs a Sume bulk run window
OpenAI lists GPT-6.1 Sol limits from 500 RPM at Tier 1 to 15,000 RPM at Tier 5. A Sume bulk run uses a concurrency window of 1-16 and 100 items.

OpenAI's model page lists GPT-6.1 Sol rate limits by usage tier, from 500 requests per minute and 500,000 tokens per minute at Tier 1 to 15,000 requests per minute and 40,000,000 tokens per minute at Tier 5. Those limits belong to an OpenAI account calling the OpenAI API. A Sume bulk run has its own controls: a concurrency window of 1 to 16 and between 1 and 100 items.
The two do not describe the same thing, and this post does not claim how one maps to the other. It sets out each so you can plan around the one you actually call.
What are the vendor limits?
The model page's rate-limit table gives requests per minute (RPM) and tokens per minute (TPM) for each tier.
| Tier | RPM | TPM |
|---|---|---|
| Tier 1 | 500 | 500,000 |
| Tier 2 | 5,000 | 1,000,000 |
| Tier 3 | 5,000 | 2,000,000 |
| Tier 4 | 10,000 | 4,000,000 |
| Tier 5 | 15,000 | 40,000,000 |
What does a Sume bulk run limit?
The bulk runs doc defines a queue of up to 100 Format runs. concurrency is a required integer from 1 to 16 and says how many child runs stay in flight at once; items is a required array of 1 to 100 entries, each the same body as a single run. Because each item is a full run body, each can carry its own model.
The doc says workspace generation concurrency still applies on top of the window, and that a 429 rate_limited means wait retry-after seconds. Creating the queue spends the write budget; polling spends a separate, larger read budget, so a poll loop does not starve your creates.
Do the vendor tiers apply to my Sume run?
The Sume docs do not say. The vendor table is the limit of an OpenAI account; the Sume docs describe the limits of the Sume API. Read the 429 behavior in the call doc and plan to Sume's retry-after, not to the vendor's RPM. If you also call OpenAI directly in the same pipeline, the RPM and TPM columns above are the budget for that part.
A small sizing example
Say you plan 60 Format runs and each takes about two minutes. With concurrency 6, ten rounds of six runs would take roughly twenty minutes; with 16 it is under four rounds, so about eight minutes. The numbers are an illustration of the window, not a measured run time. The real time depends on the Formats and on workspace generation concurrency.
Raise the window only as far as your workspace allows and back off on a 429.
What does the vendor page add beyond rate limits?
The same OpenAI page lists a 1,050,000-token context window, a 128,000-token maximum output, and says tool calling requires the Responses API. It also says the prices differ between prompts under and over 272K input tokens. For a pipeline that sends many large prompts, the TPM column often binds before the RPM column does: at Tier 1, a 500,000 TPM budget is exhausted by two requests of 250,000 tokens in a minute, long before 500 requests are made.
That arithmetic applies to a direct OpenAI integration. For Sume, the unit to plan around is the Format run and its receipt.
How should a pipeline handle both?
Treat the two limit systems as independent until a doc says otherwise.
- Use the Sume
concurrencywindow to bound Format runs in flight; start low and raise it. - Honor
retry-afteron any 429 from the Sume API, and use a separate back-off for any direct OpenAI calls. - Poll the queue at a modest interval; polling has its own read budget.
- Record the echoed
modeland usage fields from each receipt so cost and model changes are visible.
Sources
Related posts
More in Developers
- GPT Image 2.5 cost per image by quality: the token math on Sume
Sume's docs: GPT Image 2.5 output is $30 per 1M tokens. At 1024x1024, xhigh is $0.09366 and max is $0.21072, before input tokens and Sume pricing.
- GPT Image 2.5 curl command: generate and download in a shell
A copy-paste curl call to Sume's POST /v1/images for GPT Image 2.5, with jq to pull the URL, download the file, and a check for the 202 job response.
- GPT Image 2.5 custom size for a 4:5 post: why 1080x1350 is rejected
Custom image_size on Sume's gpt-image-2.5 needs both edges as multiples of 16. 1080 and 1350 are not, so use aspect_ratio 4:5 or 1088x1360 and crop.
- Use a local photo as a GPT Image 2.5 reference: it needs a URL
Sume's input_references take public HTTPS image URLs only; localhost and private URLs are rejected. Three ways to turn a file on disk into a usable reference.
Written by Sume