Synthesia Express-3 follows script sentiment; Sume's script controls
Synthesia Express-3 avatars follow script sentiment. Sume Avatar Video is steered by script text, silence beats, scene direction and quality tiers.
Synthesia's Express-3 is an avatar model whose avatars follow the sentiment of your script, so they can look happy, serious, neutral or friendly depending on what you wrote. Sume's docs describe no sentiment setting; on Sume you direct a talking video with the script text, scene direction, silence beats and a quality tier.
Express-3 facts are from Synthesia's launch post. Sume facts are from Generate avatar video, read 2026-10-01.
What does Express-3 change?
The post calls it Synthesia's most advanced avatar model for lip sync, natural body movement and gestures, with more script-aligned gestures so avatars feel less static. It says Express-3 is rolling out in stages, starting with new stock avatars and expanding across eligible avatar experiences over time.
What can I control in a Sume avatar video?
Send exactly one of script or video_inputs to POST /v1/avatar-1.0/talking-video. Total planned duration must estimate at 4 to 60 seconds.
| Control | How it works |
|---|---|
| Wording | The spoken text in script or per-scene voice |
| Silence beat | voice.type: "silence"; duration required, no script |
| Scene look | scene prompt or photo, or scene background |
quality: "plus" | Default; balanced quality path |
quality: "max" | Highest quality tier; slower turnaround |
How do I get a different mood from the same avatar?
Change the wording and the scene direction, then compare outputs. The docs make no claim that the avatar's expression is derived from the sentiment of the text, so treat mood as something you test, not something you can assume.
Where does max fit?
The docs name max as the highest tier with slower turnaround and standard as the fastest path; pick per job. Preview stills are tier-independent, so review composition first; related reading is Synthesia vs Runway.
Sources
Related posts
More in Models
- Synthesia Interactive Avatar API vs Sume rendered avatar clips
Synthesia headlines a live Interactive Avatar API. Sume avatar video is script-driven and rendered as a job: submit, poll, then fetch the clip.
- Tavus 50 new Phoenix-4.5 stock faces vs a Sume avatar handle
Tavus added 50 first-party Phoenix-4.5 stock faces, 28 of them Pro. On Sume you create your own avatar from a prompt, profile or image and reuse its handle.
- VEED lipsync-v2 video + audio vs Sume still + audio lip sync
fal lists veed/lipsync-v2 as video plus audio in. Sume lip sync routes start from a still and an audio clip, so existing footage is not re-synced.
- Veo 3.1 only makes 16:9 and 9:16; which Sume video models add more
Google's Veo 3.1 supports 16:9 and 9:16. On Sume, Seedance, MiniMax and Wan add 4:3, 1:1 and 3:4, Kling adds 1:1, and Grok takes no ratio.
Written by Sume