HeyGen burst concurrency bills 1.5x; Sume queues avatar jobs
HeyGen Enterprise burst adds up to 50 slots at 1.5x the contract rate. Sume queues valid avatar jobs by plan instead; here is how admission works.
HeyGen's September 2026 burst concurrency lets Enterprise workspaces run extra renders at 1.5 times the contract rate. Sume does not sell a burst tier for avatar videos: a valid job is accepted as queued while queue capacity remains, and it starts when a plan slot opens.
The practical difference is who pays for waiting. With a burst surcharge you pay more to skip the line. With Sume's queue-first admission you wait, at the same price, unless the queue itself is full.
What did HeyGen actually ship?
According to HeyGen's API changelog, Enterprise workspaces can switch on burst capacity from Capacity, then Concurrency settings. It adds up to 50 extra concurrent slots, set separately for avatar renders and for video translations. Workflows admitted while every included slot is busy bill at 1.5 times the contract rate, and they are tagged Burst in the Activity tab.
That is an Enterprise feature described on one changelog entry. The changelog does not say what other plans do when they hit their limit, so this post makes no claim about that.
How does Sume admit an avatar video when slots are busy?
Sume separates concurrency (jobs in processing) from queue capacity (jobs accepted but waiting). Concurrency being full is not an error by itself. A valid POST /v1/avatar-1.0/talking-video returns a job id and sits in queued until a worker moves it to processing under your workspace's concurrency guard.
Concurrency is plan-only: prepaid top-ups do not raise it. The default queue capacity is max(3, concurrency_limit x 5).
| Plan | Processing concurrency | Queue capacity | Accepted jobs |
|---|---|---|---|
| Free | 1 | 5 | 6 |
| Pro | 4 | 20 | 24 |
| Startup | 8 | 40 | 48 |
| Scale | 20 | 100 | 120 |
| Enterprise | 20 | 100 | 120 |
What happens when the queue is full?
Only then does the submit fail, with 429 queue_full. Sume releases or refunds the reservation for the failed admission. The docs tell clients to stop adding work, poll existing jobs until one finishes, cancel queued jobs they no longer need, and retry with the same Idempotency-Key.
Cancellation works only before generation starts. After that you get 409 job_generation_already_started and the job finishes normally.
429 rate_limitedis request volume, not capacity: back off usingretry-afterwhen present.402 insufficient_creditsmeans the estimated cost could not be reserved from your balance.- Sume exposes queue counts, not a per-job queue position or an ETA.
Can I pay Sume to run more at once?
Not through top-ups. The docs state that generation concurrency is plan-only, and that admin overrides can raise the effective concurrency_limit (for example for Enterprise contracts). The dashboard Concurrency tab shows the effective value, and submit responses carry it as generation_limits.concurrency_limit.
The docs we read describe no surcharge multiplier for jobs that wait. If you need a hard ceiling on turnaround, plan the batch around the effective limit rather than around a premium lane.
How should I size a batch of avatar videos?
Use max(0, concurrency_limit - active_generation_jobs - queued_generation_jobs), capped by queue_capacity_remaining, as your budget for new in-flight work. The wave_size_hint field is only a submission-wave hint, not a processing width.
Here is a minimal submit that is safe to retry. The same key on the same body returns the same job instead of billing twice.
curl -X POST https://api.sume.com/v1/avatar-1.0/talking-video \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: avatar-batch-007-item-012" \
-d '{"avatar_handle":"sume_clawra","script":"Welcome back. Here is this week in 20 seconds.","quality":"standard","mode":"async"}'How do I watch a queued job?
Store the job_id from the first response, then poll GET /v1/jobs/{id}/status with backoff and honour next_poll_after_seconds when it is present. Stop when terminal is true, and fetch the result once result_ready is true. A client-side timeout does not cancel the job; it keeps running and still bills, so a worker that gives up should resume from the stored status_url instead of resubmitting.
Webhook mode is the quieter option for batches: it delivers only job.completed, job.failed and job.canceled, and the docs recommend keeping status polling as a backup for deliveries that never arrive.
Read the Generation admission page for the full field list, and the Generate avatar video page for the 4-60 second window every script must estimate into. For the pacing arithmetic in more depth, see pace bulk submits with generation_limits.
Sources
Related posts
More in Comparisons
- HeyGen free plan limits vs paying per second on Sume
HeyGen's free plan lists 3 videos a month of up to 1 minute. Here is what that means for a test project, and how Sume's per-second billing compares.
- HeyGen photo avatar renders without a reference look (Aug 2026)
HeyGen's August 2026 change lets an Avatar V photo avatar render from the photo alone; motion_prompt still needs a reference look. Sume's photo path, compared.
- HeyGen pause tag vs a Sume silence scene in an avatar video
HeyGen supports break tags up to 5 s on professional voice clones. Sume has no inline pause tag documented; you add a silence scene inside video_inputs instead.
- HeyGen Studio Templates API vs a reusable Sume avatar request
HeyGen can now create templates from videos with POST /v3/templates. Sume has no avatar-video template object; the request body plus a handle is the template.
Written by Sume