HeyGen burst concurrency bills 1.5x; Sume queues avatar jobs

HeyGen Enterprise burst adds up to 50 slots at 1.5x the contract rate. Sume queues valid avatar jobs by plan instead; here is how admission works.

5 min readSume
All posts

HeyGen's September 2026 burst concurrency lets Enterprise workspaces run extra renders at 1.5 times the contract rate. Sume does not sell a burst tier for avatar videos: a valid job is accepted as queued while queue capacity remains, and it starts when a plan slot opens.

The practical difference is who pays for waiting. With a burst surcharge you pay more to skip the line. With Sume's queue-first admission you wait, at the same price, unless the queue itself is full.

What did HeyGen actually ship?

According to HeyGen's API changelog, Enterprise workspaces can switch on burst capacity from Capacity, then Concurrency settings. It adds up to 50 extra concurrent slots, set separately for avatar renders and for video translations. Workflows admitted while every included slot is busy bill at 1.5 times the contract rate, and they are tagged Burst in the Activity tab.

That is an Enterprise feature described on one changelog entry. The changelog does not say what other plans do when they hit their limit, so this post makes no claim about that.

How does Sume admit an avatar video when slots are busy?

Sume separates concurrency (jobs in processing) from queue capacity (jobs accepted but waiting). Concurrency being full is not an error by itself. A valid POST /v1/avatar-1.0/talking-video returns a job id and sits in queued until a worker moves it to processing under your workspace's concurrency guard.

Concurrency is plan-only: prepaid top-ups do not raise it. The default queue capacity is max(3, concurrency_limit x 5).

Sume plan concurrency and queue defaults (Generation admission docs, read 2026-10-02)
PlanProcessing concurrencyQueue capacityAccepted jobs
Free156
Pro42024
Startup84048
Scale20100120
Enterprise20100120

What happens when the queue is full?

Only then does the submit fail, with 429 queue_full. Sume releases or refunds the reservation for the failed admission. The docs tell clients to stop adding work, poll existing jobs until one finishes, cancel queued jobs they no longer need, and retry with the same Idempotency-Key.

Cancellation works only before generation starts. After that you get 409 job_generation_already_started and the job finishes normally.

  • 429 rate_limited is request volume, not capacity: back off using retry-after when present.
  • 402 insufficient_credits means the estimated cost could not be reserved from your balance.
  • Sume exposes queue counts, not a per-job queue position or an ETA.

Can I pay Sume to run more at once?

Not through top-ups. The docs state that generation concurrency is plan-only, and that admin overrides can raise the effective concurrency_limit (for example for Enterprise contracts). The dashboard Concurrency tab shows the effective value, and submit responses carry it as generation_limits.concurrency_limit.

The docs we read describe no surcharge multiplier for jobs that wait. If you need a hard ceiling on turnaround, plan the batch around the effective limit rather than around a premium lane.

How should I size a batch of avatar videos?

Use max(0, concurrency_limit - active_generation_jobs - queued_generation_jobs), capped by queue_capacity_remaining, as your budget for new in-flight work. The wave_size_hint field is only a submission-wave hint, not a processing width.

Here is a minimal submit that is safe to retry. The same key on the same body returns the same job instead of billing twice.

curl -X POST https://api.sume.com/v1/avatar-1.0/talking-video \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: avatar-batch-007-item-012" \
  -d '{"avatar_handle":"sume_clawra","script":"Welcome back. Here is this week in 20 seconds.","quality":"standard","mode":"async"}'

How do I watch a queued job?

Store the job_id from the first response, then poll GET /v1/jobs/{id}/status with backoff and honour next_poll_after_seconds when it is present. Stop when terminal is true, and fetch the result once result_ready is true. A client-side timeout does not cancel the job; it keeps running and still bills, so a worker that gives up should resume from the stored status_url instead of resubmitting.

Webhook mode is the quieter option for batches: it delivers only job.completed, job.failed and job.canceled, and the docs recommend keeping status polling as a backup for deliveries that never arrive.

Read the Generation admission page for the full field list, and the Generate avatar video page for the 4-60 second window every script must estimate into. For the pacing arithmetic in more depth, see pace bulk submits with generation_limits.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume