30 writes a second is 1,800 a minute: sizing Sume plans against it

HeyGen's enterprise write limit is 30 per second. On Sume that is 2 to 20 submits per second by plan, and accepted-job capacity binds long before it.

4 min readSume
All posts

A per-second limit and a per-minute limit are easy to compare once you convert. HeyGen's changelog (read 2026-10-07) says enterprise write operations rose to 30 per second, from 10. That is 1,800 writes a minute. Sume's write budgets are per minute by plan: 120 on Free, 300 on Pro, 600 on Startup and 1,200 on Scale, which is 2, 5, 10 and 20 submits per second if you spread them evenly.

But request rate is rarely what stops a Sume batch. The number of accepted generation jobs is far smaller than the minute budget on every plan, so you hit queue_full long before a rate limit.

The conversion

Per-second figures below are the plan's per-minute write budget divided by 60, rounded to one decimal. Accepted capacity is concurrency plus queue, from the generation admission page.

Sume writes per minute vs accepted jobs, by plan (read 2026-10-07)
PlanWrites per minutePer secondAccepted jobs (processing + queued)
Free1202.06
Pro3005.024
Startup60010.048
Scale120020.0120

Reading the table

On Scale you may send 1,200 writes a minute, but only 120 paid generation jobs can be processing or queued at once. Submit 150 at a burst and 30 get 429 queue_full, not rate_limited. The two codes need different handling: rate_limited carries retry-after; queue_full means wait for a job to finish or cancel one, then retry with the same idempotency key.

Cancels, uploads and run creation are also writes, so a client that cancels aggressively spends the same budget as one that submits.

Size a burst before sending it

Use generation_limits from submit responses rather than the static table: concurrency_limit is the effective processing cap and queue_capacity_remaining says how many more jobs fit. The docs' own guidance is to size new in-flight work as max(0, concurrency_limit - active - queued), limited to queue_capacity_remaining, and to treat wave_size_hint as a submission hint only.

HeyGen's 30 per second is HeyGen's contract tier, not a statement about Sume. Enterprise on Sume is not self-serve: until a contracted number is provisioned an Enterprise key uses the Scale row.

What to do with this on Monday

Do not size a worker pool from the headline rate. Size it from accepted capacity, and let the rate limit be the safety net.

If you run several services against one workspace, remember that generation_limits describes the workspace, not your process: another service's queued jobs shrink your headroom. Read the snapshot on every submit response and treat it as a snapshot that can change a moment later.

  • Submit in waves no larger than your computed headroom.
  • Poll with the reads budget, which is forty times the writes number, and stop polling jobs that are terminal.
  • On queue_full, stop adding work, cancel queued jobs you no longer need, and retry with the same key after capacity opens.
  • Re-read your plan's row in the dashboard Concurrency tab; the docs say to prefer the effective field to any static table.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume