Together AI dynamic rate limits (429, 503) vs Sume rate_limited
Together AI publishes no fixed tiers: limits track live capacity and your recent traffic. How its 429 and 503 map to Sume's rate_limited and queue_full.

Together AI does not publish a fixed rate-limit table: its serverless limits are dynamic per model, so you plan around the response, not a number. Sume answers the same problem with two named 429 codes, rate_limited and queue_full, plus ratelimit-* and retry-after headers, so a retry loop can tell a busy window from a full workspace.
Together's side comes from its Rate Limits page and Sume's from Errors and rate limits and Generation admission, all read on 2026-10-02. Nothing below is a benchmark; it is a map between two error vocabularies.
How does Together AI decide your rate limit?
Together's page says serverless limits are dynamic per model and adjust to the model's live capacity and your recent successful usage. It adds that sustained, successful traffic raises your dynamic rate over time, while sudden spikes far beyond recent usage may be throttled.
The consequence for a video or image batch is that day one and day thirty can differ. A job that ramps from a few requests to hundreds in one minute is the pattern the page warns about; a steady ramp is the pattern it rewards. If you need fixed, known limits for planning, the page points to dedicated endpoints instead of serverless models.
What do Together's 429 and 503 mean?
A 429 carries one of two error types: dynamic_request_limited for request-based limiting and dynamic_token_limited for token-based limiting. The x-ratelimit-reset header appears on 429 responses and gives the suggested retry interval for the model. A 503 is different: the request was at or below your dynamic rate, but the platform was overloaded.
Together also notes that successful requests come back without rate-limit headers, so the presence of x-ratelimit-reset is itself the throttle signal.
How does that map to Sume's errors?
Sume has no per-model dynamic rate and no token-based limiter in its docs. It splits the problem differently, as the table shows.
| Situation | Together AI | Sume |
|---|---|---|
| Too many requests | 429 dynamic_request_limited | 429 rate_limited |
| Token-based throttle | 429 dynamic_token_limited | No equivalent documented |
| Retry hint | x-ratelimit-reset on 429 | retry-after when present, plus ratelimit-limit, ratelimit-remaining, ratelimit-reset |
| Platform overloaded | 503 at or below your rate | 503 provider_capacity_exceeded; retry later with the same idempotency key |
| Workspace full of paid jobs | Not a separate code | 429 queue_full |
Why does queue_full matter for batches?
Sume says queue_full is different from ordinary request rate limiting: the workspace has no room to accept another paid generation job until an existing queued or processing job finishes or is canceled. Concurrency being full by itself is not an error. Sume accepts valid jobs as queued while queue capacity remains.
So a Sume batch has three outcomes where a Together batch has two: accepted and running, accepted and queued, or rejected with queue_full. Cancel queued jobs you no longer need, then resubmit with the same idempotency key. Sume's docs say not to retry unsafe submit requests without an Idempotency-Key.
What should a port do with the retry loop?
Replace the check for x-ratelimit-reset with a check on the status code and error code. On rate_limited, honor retry-after when it is present and back off otherwise. On queue_full, stop submitting and wait for jobs to finish. On provider_capacity_exceeded, retry later with the same key. Sume does not publish numeric request limits on these pages, so do not hard-code one.
Sume does not offer dedicated endpoints with fixed throughput on these pages either. If you need a contracted rate, ask Sume rather than assuming one.
Sources
Related posts
More in Comparisons
- Together AI says download videos immediately: Sume media URLs
Together's docs say not to rely on video URLs for long-term storage. Sume returns media.sume.com artifacts as its public contract. What to save in each case.
- Together AI video status values vs Sume job statuses (canceled)
Together's video job has five statuses, including cancelled. Sume has five too, but spells it canceled and adds terminal and result_ready flags. Map them here.
- Veo 3.1 1080p 8-second clip cost vs Sume per-second rows
What an 8-second 1080p clip costs on Veo 3.1 Standard, Fast and Lite list prices versus five 1080p rows in the Sume video catalog, dated 2026-10-01.
- Veo 3.1 4K price per second vs Gemini Omni Flash 4K
Veo 3.1 4K costs $0.60 per second on Standard and $0.30 on Fast. Gemini Omni Flash 1.1 on Sume lists 4K at $0.375 per second. Dated 2026-10-01.
Written by Sume