Synthesia API documentation: the quickstart, step by step
Synthesia's API docs make a video in four steps: a Legacy (v2) key, POST /v2/videos, poll until complete, then download from a time-limited link.

Synthesia's API documentation starts with a four-step Video API quickstart: create an API key with the Legacy (v2) scope, send POST /v2/videos with a script, an avatar and a background, poll the video until its status is complete, then download the MP4 from a time-limited download link. A webhook can replace the polling.
Every Synthesia detail below comes from its quickstart, its Create a video reference and its webhook pages, read on 2026-09-28. The Sume section at the end comes from Generate avatar video.
How do I get a Synthesia API key?
In Synthesia, select Developers in the sidebar, open the API keys tab, choose New API key, pick the Legacy (v2) scope and an expiration, and copy the key: Synthesia shows it only once. An Interactive Avatars-scoped key won't work for video requests.
The key goes in the Authorization header of each request. Synthesia notes that a key belongs to your account, not the workspace, so any webhook you create with it is tied to you. Plans, credits and rate limits are covered in Synthesia API pricing.
What does the create-video request look like?
The quickstart lists the values as a flat table, but in the endpoint's schema the script, avatar and background sit inside input, an array where each object is one clip (scene) of the video. input is the only required top-level field; each clip requires avatar and background.
| Field | What it does |
|---|---|
test | true makes a free test video with a watermark, not counted toward your quota; 30 test videos per day |
title | Title shown on the video's share page |
input[].scriptText | The script the avatar speaks, in any of Synthesia's supported languages |
input[].avatar | A stock or custom avatar id, such as anna_costume1_cameraA |
input[].background | A stock background such as green_screen, an uploaded asset id, or a URL |
aspectRatio | 16:9 (default), 9:16, 1:1, 4:5 or 5:4 |
visibility | private (default) or public; private videos download only through the time-limited link |
curl -X POST https://api.synthesia.io/v2/videos \
-H "Authorization: $SYNTHESIA_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"test": true,
"title": "My first Synthetic video",
"input": [{
"scriptText": "Hello, World! This is my first synthetic video.",
"avatar": "anna_costume1_cameraA",
"background": "green_screen"
}]
}'How do I know when the video is ready?
Store the video id from the response. Synthesia starts processing at once, and its quickstart says this can take 3 to 5 minutes.
- Poll
GET /v2/videos/{video_id}(Retrieve a video). statusis one ofin_progress,complete,error,rejected,deletedorapproved.- At
complete, the response carries a time-limited link indownload. Fetch it right away and keep your own copy of the.mp4.
Can I use a webhook instead of polling?
Yes. Create a Webhook (POST /v2/webhooks) takes a url and the events to send: video.completed, video.failed or both.
- The response includes a
secret, which is only available when the webhook is created. Save it then. - Synthesia expects a non-error response within six seconds; if delivery fails, it retries twice over a ten-minute window.
- Each event carries
Synthesia-TimestampandSynthesia-Signatureheaders. The signature is an HMAC SHA-256 of the timestamp and the raw body joined with., keyed with the webhook secret. Synthesia's sample compares with==; a constant-time comparison is safer. As an optional step, Synthesia suggests rejecting old timestamps.
What does the same job look like on Sume's API?
On Sume, the script-to-avatar job is one POST /v1/avatar-1.0/talking-video with a ready avatar_handle and a script that Sume estimates at 4-60 seconds. The key goes in the header as Bearer $SUME_API_KEY. The default async mode answers 202 with the job and its polling URLs; add webhook_url and Sume also sends a signed terminal event (job.completed, job.failed or job.canceled), with polling kept as a backup (Jobs and results, Webhooks).
Completed results include public media.sume.com video artifacts. In current code the avatar speaks English only. The full request, options and rates are in Talking avatar video API, and Sume vs Synthesia compares the two products.
Sources
- Synthesia Video API Quickstart (read 2026-09-28)
- Synthesia: Create a video (read 2026-09-28)
- Synthesia Video API Introduction (read 2026-09-28)
- Synthesia: Retrieve a video (read 2026-09-28)
- Synthesia Webhooks (read 2026-09-28)
- Synthesia: Create a Webhook (read 2026-09-28)
- Synthesia: Verifying signatures (read 2026-09-28)
- Generate avatar video
- Jobs and results
- Webhooks
Related posts
More in Developers
- Text to image API: send a prompt, get image URLs back
A text-to-image API turns a prompt sent over HTTPS into generated images. How to call one: the request, the response, slow jobs, and the cost.
- Text to speech API in Java: pick a voice, save the MP3
Call a text to speech API from Java: choose a voice selector and an audio format, POST the text, wait for the job, then write the MP3 to disk.
- Text to video and image to video: how they differ
Text-to-video invents every frame from words; image-to-video starts on your picture and animates it. How they differ, and how one request does both.
- Text to video API: how it works, what it returns, cost
A text to video API takes a prompt and returns a job id, not a video: poll it or take a webhook, then download the file. Fields, flow and prices.
Written by Sume