Synthesia Survey Sessions vs Sume: not interactive, but scripted
Synthesia's Survey Sessions avatar asks questions live and reports themes. Sume renders avatar question clips from a script, with no live conversation.

Sume does not run live avatar conversations. It renders an avatar video from a script, one clip at a time. Synthesia's Sessions page, read 2026-10-01, says Survey Sessions are live: the avatar asks questions, follows up and produces an AI report of themes and sentiment, plus transcript, audio and CSV.
What can I build on Sume?
POST /v1/avatar-1.0/talking-video takes a script or video_inputs. The estimated duration must be 4 to 60 seconds, there is one avatar per final video, quality is standard, plus (default) or max, and resolution is 720p. Render each survey question as its own clip.
How do I collect answers?
Sume does not collect them. Play the clip in your own form or app, record the reply there, and send the recording to POST /v1/stt-1.0/transcribe to get text and word timings. The follow-up logic and the theme report would be yours.
What does this cost you in features?
No live follow-up questions, no sentiment report, no CSV export from Sume. You gain a fixed per-clip render and your own data ownership. If live interviews are the point, Synthesia's feature is the fit.
| Feature | Synthesia Survey Sessions | Sume |
|---|---|---|
| Live questions and follow-ups | Yes | No, one scripted clip at a time |
| AI report of themes and sentiment | Yes | No |
| CSV export | Yes | No |
| Clip length | Not listed here | 4 to 60 seconds |
Sources
Related posts
More in Sume Avatar 1.0
- VEED Fabric Emotions audio tags vs Sume Fabric audio_url
VEED's Fabric Emotions adds inline tags like [excited] in the script. On Sume, Fabric takes a finished audio_url; emotion lives in the TTS step.
- Introducing Sume Avatar 1.0
Sume Avatar 1.0 is a multi-agent orchestration system as a single avatar model.
- Avatar Face Swap API (Beta): apply an avatar face to a video
Avatar Face Swap 1.0 is a Beta Sume endpoint that applies a ready avatar's face to a short public source video. Required fields, limits, and polling.
- Avatar video previews: approve the first frame before rendering
Create an avatar video preview to get first-frame stills, regenerate them if needed, then call generate-video on the preview id to render the final video.
Written by Sume