Synthesia Interactive Avatar and LiveKit: Sume's job-based pieces

Synthesia's Interactive Avatar API runs on a LiveKit plugin with your own LLM and STT. Sume offers job pieces: TTS, speech to text and still-plus-audio clips.

4 min readSume
All posts

Synthesia's Interactive Avatar API is a live system: you supply the LLM, speech-to-text and optionally text to speech, and Synthesia layers a lip-synced avatar on top through its livekit-plugins-synthesia LiveKit plugin. Sume has no live session, but it covers some of the surrounding pieces as jobs: speech to text, text to speech and a still-plus-audio talking clip.

Synthesia facts are from its Interactive Avatar post, which lists the API as launched for Enterprise customers, with a managed full-stack pipeline listed for October 1, 2026. Sume facts are from Models and Jobs and results, read 2026-10-01.

Which parts of that stack could run on Sume?

Only the batch parts. The table maps each layer to what the Sume docs describe.

Layers of an avatar stack and Sume equivalents, docs read 2026-10-01.
LayerOn Sume
LLMNot part of the documented jobs listed here
Speech to textPOST /v1/stt-1.0/transcribe, public id sume/stt-1.0
Text to speechPOST /v1/tts-1.0/generate or the TTS Router
Talking clipveed/fabric-1.0: talking still plus audio
Real-time avatar sessionNot documented

How does a talking clip get made?

Generate speech, then send the audio (a Sume-hosted URL) and a still to POST /v1/veed/fabric-1.0 with audio_url, measured duration_seconds and one visual source. The models page says video models do not lip-sync to generated TTS or to a later voice-over, so a talking face goes through Fabric with an accepted still plus TTS.

How should I wait on those jobs?

Prefer async submit with an idempotency key. wait_timeout_seconds is clamped to 0..30 and bounds only the HTTP wait. Poll until the job is terminal, or take a webhook; see Generation admission.

When is this the wrong fit?

When a user talks and the avatar must reply in real time. Use a live system for that and Sume for rendered clips; compare live conversation vs rendered clips.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume