Talking-video API status values: HeyGen, Synthesia, D-ID, Sume mapped
Four avatar-video APIs, four status vocabularies: pending, in_progress, started, queued. Map them to one state machine in your code (read 2026-10-10).

If you call more than one talking-video API, map every vendor status into four states of your own: waiting, running, done and stopped. HeyGen uses pending, processing, completed and failed; Synthesia lists in_progress, complete, error, rejected, approved and deleted; D-ID lists created, started, done, error and rejected; Sume uses queued, processing, completed, failed and canceled.
The vendor lists come from HeyGen's Digital Twin page, Synthesia's Create a video, and D-ID's Create a talk and Create a Video, all read 2026-10-10. Sume's is from Jobs and results.
The mapping
Where a vendor page does not say whether a value is terminal, the cell says what the name suggests and flags it.
| Your state | HeyGen | Synthesia | D-ID | Sume |
|---|---|---|---|---|
| Waiting | pending | Not separate on the page read | created | queued |
| Running | processing | in_progress | started | processing |
| Done | completed | complete | done | completed |
| Stopped with an error | failed | error | error | failed |
| Stopped by a rule or a person | Not listed | rejected, deleted | rejected | canceled |
| Other | None | approved | None | Status endpoint also returns IN_QUEUE / IN_PROGRESS / COMPLETED / FAILED / CANCELED |
Details that bite
Synthesia's approved and D-ID's rejected look like review states, and the pages I read do not explain them in a way I can map with confidence. Do not guess: until you have confirmed what triggers them, route them to a human-readable bucket rather than to success or failure. The D-ID page lists rejected as a creation-response value, which suggests a request refused up front, but check the reference before you rely on that reading.
On Sume, queued is a normal accepted state, not a stall. Workspace concurrency applies when workers move jobs from queued to processing. The status endpoint returns booleans, terminal and result_ready, so you can poll on those instead of matching strings, and its queue-shaped status field maps one-to-one onto sume_status.
A small adapter
Keep the vendor string in your row for debugging, but let your own code branch on the four states. A table-driven adapter per vendor is a dozen lines, and it means a new vendor value fails loudly as unknown instead of silently being treated as done.
For Sume, add one rule: a completed job is the only state where GET /v1/jobs/{id}/result is valid. For failed and canceled jobs the result route answers 409 job_not_completed, so read the failure from the job record instead.
Sume avatar-video specifics
Avatar videos have a resource as well as a job. GET /v1/avatar-videos takes a status filter with queued, processing, completed, failed, canceled and ready, where ready is an alias for completed jobs, and a limit of 1 to 100. With previews there are two jobs, one for the first-frame stills and one for the final render, and the preview resource reports resource_status and job_status, which the docs prefer over the legacy status field. See Avatar video previews.
Cancellation is only possible before generation starts. After that the API returns 409 job_generation_already_started with details.cancelable: false, so a render that has started is a render you will pay for. If you need the option to stop, use the preview stage: approve the first frame, then start the final render.
What to do on each state
Waiting and running states deserve a poll or a webhook wait, nothing more. Done states should trigger one fetch of the output and one write to your own store. Stopped states should carry a reason where the vendor gives one, and should never be retried blindly: a refused script that is resubmitted unchanged will be refused again.
On Sume, a retry after a timeout on your side is safe if you resend the same Idempotency-Key, which returns the original job rather than starting a second paid render. A deliberate re-run after a failure should use a new key, so it is a new job on purpose.
Sources
Related posts
More in Comparisons
- Tavus lists 42 languages; Sume Avatar 1.0 speaks English only
Tavus video generation lists 42 languages. Sume Avatar 1.0 speaks English only. A short decision guide based on who is listening and what you can accept.
- Tavus callback_url payloads: what to check before you trust one
Tavus's webhooks page lists conversation events but no signature or retry rule. How to treat them, vs Sume's signed events (read 2026-10-10).
- Text-only hero art: Imagen 4 Ultra or Nano Banana 2.1 on Sume
Imagen 4 Ultra costs 0.075 USD, takes no references and lists 5 ratios. Nano Banana 2.1 costs 0.10 and lists 15. Which to pick for 100 hero images.
- TikTok ad profile photo is 98x98 under 50 KB: what Sume cannot output
TikTok in-feed ads want a 98x98 px profile photo under 50 KB. No Sume route outputs that size, so generate the picture and resize it outside Sume.
Written by Sume