HeyGen enterprise writes rise to 30/s; how Sume limits submits

HeyGen raised enterprise limits: writes 30 per second, polling 1,500 per minute. Sume publishes no per-second number; it uses plan queues and rate headers.

4 min readSume
All posts

HeyGen's API changelog now lists higher Enterprise rate limits: reads 1,000 per minute, writes 30 per second, heavy calls 120 per minute and status polling 1,500 per minute. Sume does not publish a requests-per-second figure for avatar submits; what its docs give you is plan-based processing concurrency, queue capacity, and rate-limit headers.

The HeyGen numbers are from HeyGen's API changelog, read 2026-10-11. The entry says the old values were 300 reads per minute, 10 writes per second, 30 heavy calls per minute and 500 status polls per minute, and that Pay-As-You-Go and subscription limits are unchanged. Sume's side is from Generation admission and Errors and rate limits.

What HeyGen's change says

The change only applies to Enterprise workspaces. The page does not give the standard-plan numbers, so this post cannot compare a small account. If you are on HeyGen's lower plans, nothing changed for you.

HeyGen Enterprise limits before and after (changelog read 2026-10-11)
Call classBeforeAfter
Read300 per minute1,000 per minute
Write10 per second30 per second
Heavy30 per minute120 per minute
Status polling500 per minute1,500 per minute

How Sume separates the controls

Sume's docs split four controls that people tend to mix up. Processing concurrency is how many paid jobs run at once. Queue capacity is how many accepted jobs can wait. Submit rate limits are request volume on submit endpoints. Balance is the last: a request fails with 402 when the reservation cannot be covered.

Concurrency is a dispatch limit, not a submit limit. A Free workspace processes 1 job at a time with a queue of 5, so 6 accepted jobs; Pro processes 4 with a queue of 20 (24 accepted); Startup 8 and 40 (48); Scale and Enterprise 20 and 100 (120). The default queue is the larger of 3 or five times the concurrency, and the dashboard and generation_limits field show the effective values.

What that means for an avatar batch

A batch of avatar videos does not hit a per-second wall on Sume. It hits the accepted-job ceiling: submit more than the plan allows and the next paid generation returns 429 queue_full, which is different from 429 rate_limited. Wait for jobs to finish or cancel queued ones, then retry with the same idempotency key. rate_limited is the abuse-protection response; use retry-after when present, and back off.

Polling is where the two designs differ most. HeyGen publishes a status-poll budget. Sume says read, status and list endpoints can have rate limits and tells you to treat them as poll backpressure, with the response headers ratelimit-limit, ratelimit-remaining, ratelimit-reset and retry-after as the signal. A webhook avoids most polling, though the docs say to keep polling available as a backup.

A plan for 100 clips

On a Pro workspace the arithmetic is 24 accepted jobs at once, so a 100-clip batch is submitted in waves of 24, with each wave refilled as jobs finish. Give every clip a stable idempotency key so a retry after 429 never makes a duplicate, and read generation_limits from a submit response instead of hard-coding the table above.

If you came here because HeyGen's write limit rose and you want the same headroom, the equivalent move on Sume is a higher plan, since concurrency is plan-only and prepaid top-ups do not raise it.

Finally, read the error before you tune anything. A 402 means the balance cannot cover the reservation, a 429 queue_full means the accepted-job ceiling is reached, and a 429 rate_limited means request volume. Each has a different fix, and only the last is solved by slowing down. Teams that treat every 429 the same end up sleeping while a queue drains, or hammering while a limit resets.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume