Synthesia API rate limits by tier vs Sume's queue and 429s

Synthesia caps writes at 60 to 120 a minute by tier and answers 429. Sume queues accepted jobs and returns queue_full or rate_limited. Read 2026-10-10.

5 min readSume
All posts

Synthesia's API documents per-tier write limits of 60 to 120 requests a minute (600 to 3,000 a day at the top end), and it answers 429 Too Many Requests when you exceed them. Sume separates two controls: a submit rate limit that returns 429 rate_limited, and an accepted-job capacity that returns 429 queue_full only when the workspace queue is also full. Concurrency alone is not an error on Sume.

Synthesia's numbers are from its Video API Introduction, read 2026-10-10, and Sume's from Generation admission.

Synthesia's tiers

The page lists write operations per endpoint, and a read range.

Synthesia write limits per endpoint (page read 2026-10-10)
TierPer minutePer hourPer day
Tier 1 (Enterprise)1206003,000
Tier 2 (Enterprise)804002,000
Tier 3 (Pro)603001,000

Reads run at 60 to 120 requests a minute and 20,000 to 80,000 a day, depending on tier. The same page limits test-mode videos to 30 a day, says the headers on a 429 tell you which limit you hit and when it resets, and states that API keys belong to the account rather than the workspace.

Sume's plan table

Sume's limit is about running jobs, not request counts. Queued is a normal accepted state.

Sume generation capacity by plan (per docs.sume.com/workflows/generation-admission; the dashboard Concurrency tab is the source of truth)
PlanProcessing concurrencyQueue capacity (default)Accepted job capacity
Free156
Pro42024
Startup84048
Scale20100120
Enterprise20100120

What a burst looks like on each

Say you want 30 personalized presenter clips tonight. On a Synthesia Pro-tier key, the minute limit is about request pacing: 30 creates fit well inside 60 a minute, so rate is unlikely to bite, and what matters is your quota and render time.

On Sume, 30 avatar videos submitted at once on a Pro workspace is 6 more than the 24 jobs it can hold as processing plus queued, so the 25th paid submit returns 429 queue_full. The fix is to submit in waves of up to the accepted capacity, store every job_id, and refill as jobs finish, retrying a refused submit with the same Idempotency-Key. Read the live numbers from the generation_limits object on a submit response rather than hard-coding the table.

A 402 insufficient_credits is a third, separate refusal: Sume cannot reserve the estimated cost from the workspace balance, and it fires before provider work starts.

Rules that carry across both

  • Treat 429 as backpressure. Sume sends retry-after on rate_limited when it can; use it.
  • Retry only with the same idempotency key on Sume, so a retried submit cannot bill twice.
  • Do not poll in a tight loop. Sume's docs say to use exponential backoff on status reads, and to obey next_poll_after_seconds when it is present.
  • Prefer a webhook plus a slow poll over a fast poll, since read endpoints have limits too.

Planning for a limit you do not control

Whichever vendor you use, put your own small queue in front of it. A queue with a fixed number of workers keeps you under any published ceiling, and lets you back off when a 429 comes back. It also makes the cost of a bulk run predictable, because you decide the pace.

A fixed-worker queue

The simplest guard is a worker pool of a fixed size, say four, pulling scripts from a list. Each worker submits one job, waits for its terminal state, then takes the next. You never exceed four in flight, a 429 only pauses one worker, and your total time is easy to estimate from the average render time.

On Sume, jobs past the workspace concurrency limit are accepted as queued rather than refused, so the pool size is a pace choice, not a protection against errors.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume