HeyGen enterprise writes rise to 30/s; how Sume limits submits
HeyGen raised enterprise limits: writes 30 per second, polling 1,500 per minute. Sume publishes no per-second number; it uses plan queues and rate headers.

HeyGen's API changelog now lists higher Enterprise rate limits: reads 1,000 per minute, writes 30 per second, heavy calls 120 per minute and status polling 1,500 per minute. Sume does not publish a requests-per-second figure for avatar submits; what its docs give you is plan-based processing concurrency, queue capacity, and rate-limit headers.
The HeyGen numbers are from HeyGen's API changelog, read 2026-10-11. The entry says the old values were 300 reads per minute, 10 writes per second, 30 heavy calls per minute and 500 status polls per minute, and that Pay-As-You-Go and subscription limits are unchanged. Sume's side is from Generation admission and Errors and rate limits.
What HeyGen's change says
The change only applies to Enterprise workspaces. The page does not give the standard-plan numbers, so this post cannot compare a small account. If you are on HeyGen's lower plans, nothing changed for you.
| Call class | Before | After |
|---|---|---|
| Read | 300 per minute | 1,000 per minute |
| Write | 10 per second | 30 per second |
| Heavy | 30 per minute | 120 per minute |
| Status polling | 500 per minute | 1,500 per minute |
How Sume separates the controls
Sume's docs split four controls that people tend to mix up. Processing concurrency is how many paid jobs run at once. Queue capacity is how many accepted jobs can wait. Submit rate limits are request volume on submit endpoints. Balance is the last: a request fails with 402 when the reservation cannot be covered.
Concurrency is a dispatch limit, not a submit limit. A Free workspace processes 1 job at a time with a queue of 5, so 6 accepted jobs; Pro processes 4 with a queue of 20 (24 accepted); Startup 8 and 40 (48); Scale and Enterprise 20 and 100 (120). The default queue is the larger of 3 or five times the concurrency, and the dashboard and generation_limits field show the effective values.
What that means for an avatar batch
A batch of avatar videos does not hit a per-second wall on Sume. It hits the accepted-job ceiling: submit more than the plan allows and the next paid generation returns 429 queue_full, which is different from 429 rate_limited. Wait for jobs to finish or cancel queued ones, then retry with the same idempotency key. rate_limited is the abuse-protection response; use retry-after when present, and back off.
Polling is where the two designs differ most. HeyGen publishes a status-poll budget. Sume says read, status and list endpoints can have rate limits and tells you to treat them as poll backpressure, with the response headers ratelimit-limit, ratelimit-remaining, ratelimit-reset and retry-after as the signal. A webhook avoids most polling, though the docs say to keep polling available as a backup.
A plan for 100 clips
On a Pro workspace the arithmetic is 24 accepted jobs at once, so a 100-clip batch is submitted in waves of 24, with each wave refilled as jobs finish. Give every clip a stable idempotency key so a retry after 429 never makes a duplicate, and read generation_limits from a submit response instead of hard-coding the table above.
If you came here because HeyGen's write limit rose and you want the same headroom, the equivalent move on Sume is a higher plan, since concurrency is plan-only and prepaid top-ups do not raise it.
Finally, read the error before you tune anything. A 402 means the balance cannot cover the reservation, a 429 queue_full means the accepted-job ceiling is reached, and a 429 rate_limited means request volume. Each has a different fix, and only the last is solved by slowing down. Teams that treat every 429 the same end up sleeping while a queue drains, or hammering while a limit resets.
Sources
Related posts
More in Comparisons
- HeyGen instant voice clone audio in Sume? Only Sume-hosted audio
HeyGen's new instant voice clone can output TTS audio, but Sume's talking-still routes take only Sume-hosted audio up to 10 MB. What it blocks and what works.
- Is AI video 'Full HD' native or upscaled? Kandinsky 6.0 vs Sume rows
Kandinsky 6.0 renders 864x480 and adds Full HD with a plug-in. Sume's H3 Max 1080p is a latent refinement from 768p. What that means for price and detail.
- Is Kandinsky 6.0 Video on Sume? No: query the catalog by need
Sume lists no Kandinsky 6.0 id. Map each Kandinsky feature to a Sume request field, then filter GET /v1/videos/models for sound and a 5 s duration.
- Kandinsky 6.0 lip-syncs in one pass; on Sume a talking face is Fabric
Kandinsky 6.0 Video makes speech and lip-sync inside the clip. Sume's docs send on-camera speech to Fabric or H3 Max Lip Sync with your audio, not a video id.
Written by Sume