Synthesia's hosted avatar pipeline (Oct 1) vs Sume clips
Synthesia's no-code hosted LLM, TTS and avatar pipeline was due Oct 1, 2026 for live conversation. If you need a rendered clip instead, Sume takes a script.
Synthesia's announcement says its full-stack pipeline, with a Synthesia-hosted LLM, text to speech and avatar and no code, would launch on October 1, 2026. That is a live, conversational avatar. If what you need is a finished video file of an avatar reading a script, that is a different product, and on Sume it is one call to POST /v1/avatar-1.0/talking-video.
This post separates the two so you can pick without reading both vendors' docs.
What did Synthesia announce?
Its Interactive Avatar API post says Interactive Avatars are available now for Enterprise customers, with a Bring-Your-Own-Stack approach: you supply your own LLM, speech-to-text and text-to-speech providers, and Synthesia renders the avatar over LiveKit. The managed full-stack option, hosting the LLM and TTS as well, was scheduled for October 1, 2026.
The page does not state pricing, rate limits or latency, so this post makes no claim about them.
Is that the same thing as a rendered avatar clip?
No. A live avatar holds a conversation: someone speaks, a model replies, the avatar answers in real time. A rendered clip is a job. You send a script, the job moves through queued, processing and a terminal state, and you fetch a public MP4 from the result.
Sume's side is the second kind. An avatar video takes a ready avatar_handle and exactly one of script or video_inputs, an estimated length of 4-60 seconds, and returns a video under media.sume.com. It does not answer a viewer's question, and nothing in this post suggests otherwise.
| Question | Synthesia Interactive Avatar | Sume avatar video |
|---|---|---|
| What it is | Real-time conversational avatar | Job-backed script-to-video render |
| Who writes the words | An LLM at run time | You, in script or video_inputs |
| Output | A live session | A public MP4 |
| Availability noted | Enterprise now; hosted pipeline due Oct 1, 2026 | Public API with an API key |
| Length | Not stated on the page | 4-60 seconds estimated, per job |
When do I want clips instead of a live session?
Pick a clip when the words are known in advance and many people will watch the same file: product explainers, onboarding steps, ad variants, a landing page welcome. A rendered file can be reviewed before it ships, which a live LLM answer cannot.
Sume adds a review step for that case. Create an avatar video preview to see first-frame stills, then call generate-video on the preview id. Changing quality at that step only changes the final render tier; changing the script or avatar needs a new preview.
What if I need both?
Many teams do. A live avatar answers questions at a kiosk or in support, while the same brand's intro, tutorials and promos are rendered once. Keep the paths separate: the rendered clips are idempotent jobs you can retry safely with an Idempotency-Key, and the live session is a runtime you operate.
Do not wait on a render with a long HTTP request. Sume's wait_timeout_seconds is clamped to 0-30, so a video job routinely outlasts it. Keep the job.id, poll status_url, and fetch result_url when the job reports result_ready or completed.
What should I verify before relying on either?
Check the date and plan. Synthesia's post describes the hosted option as scheduled for October 1, 2026 for team accounts, and Interactive Avatars as available now to Enterprise customers, so confirm on Synthesia's own pages that your plan can use it today.
On Sume, run one real job before you plan a campaign: submit with an Idempotency-Key, read the generation_limits snapshot in the response to see how many jobs your workspace can hold, and look at the preview still. These checks cost far less than discovering a plan limit with a hundred scripts queued.
Sources
Related posts
More in Sume Avatar 1.0
- Tavus Phoenix-4.5 adds cartoon and anime faces; Sume's options
Tavus Phoenix-4.5 supports cartoon, anime and Pixar-style faces from a photo or video. Sume creates avatars from a prompt, traits or a photo; test styles first.
- Two AI characters in one TikTok scene: one avatar per job, then cut
One Sume avatar job holds one avatar. For a two-person dialogue, make a job per speaker and alternate them in Timeline 1.0. Cost for 60 s of dialogue: $14.82.
- YouTube avatar vs Sume Avatar: selfie capture or prompt and photo
YouTube's avatar is made once from your own face and voice and used in its AI tools. How it is created, its limits, and how Sume's Avatar 1.0 differs.
- Introducing Sume Avatar 1.0
Sume Avatar 1.0 is a multi-agent orchestration system as a single avatar model.
Written by Sume