Synthesia Sessions alternative: avatar briefing clips by API

Synthesia launched Sessions for live avatar roleplay. Sume renders scripted avatar clips instead: what that covers in sales training, and what it does not.

5 min readSume
All posts

If you searched for a Synthesia Sessions alternative to run live sales roleplay, Sume is not one: Sume renders finished avatar videos from a script, and a rendered clip does not listen, push back, or score a trainee. What Sume does cover is the part of a training program that sits around the conversation: the scenario briefing, the "here is what good looks like" demo, and the follow-up clips. Those are scripted, and a scripted avatar clip is what Sume's Avatar 1.0 API produces.

Synthesia's own Sessions page, read on 2026-10-01, describes two products that are live: Roleplay Sessions, where an avatar plays a customer, prospect or employee and gives coaching after each attempt, and Survey Sessions, where an avatar interviews people and reports themes and sentiment. It says both can be tried free with no credit card. That is a two-way product. Sume's documentation describes a different one-way product, so the honest comparison is about which job you are doing.

Which job needs which tool

Split a roleplay program into its parts and the boundary is clear.

Synthesia column from the Sessions page, read 2026-10-01. Sume column from Generate avatar video and Models.
Part of the programSynthesia SessionsSume Avatar 1.0
Practice a live conversation with feedbackRoleplay Sessions: avatar responds, then coachesNot offered
Ask many people the same questions and collect themesSurvey Sessions: adaptive follow-up, reportsNot offered
Scenario briefing video before the roleplayPossible, in its video editorScript to talking-head clip, 4 to 60 seconds per job
Model answer demo for the manager to sharePossibleSame endpoint, a second script
Generate 50 variants of a briefing by team or regionNot covered on the Sessions pageOne API call per variant with an Idempotency-Key

Briefing clips with the Avatar 1.0 API

Send a ready avatar handle and a script to POST /v1/avatar-1.0/talking-video. Sume accepts a script when it estimates 4 to 60 seconds of speech, so a briefing longer than a minute becomes two or three jobs. Quality is standard, plus (the default) or max, the aspect ratio defaults to 9:16, and resolution is 720p. A scenario with a pause where the trainee is supposed to think can use a multi-scene video_inputs plan with a silence beat, which is a non-speaking scene with a required duration.

curl -X POST https://api.sume.com/v1/avatar-1.0/talking-video \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: briefing-renewal-call-001" \
  -d '{
    "avatar_handle": "sales_coach",
    "script": "You are about to call a customer who has not logged in for 30 days. Open with their last project, not with a discount.",
    "quality": "standard",
    "aspect_ratio": "16:9"
  }'

Limits to plan around

The practical split: run the conversation practice where the product is built for it, and use the API for the volume of scripted video around it.

  • No interaction. The avatar in a Sume clip says exactly what the script says, every time, and never hears the trainee.
  • One resolved avatar per final video. A two-person scenario is two clips, or a different tool.
  • Disclosure still applies. If trainees could mistake the clip for a live call, label it as a recording.
  • Synthesia's pricing, language count and plan limits for Sessions are on its own pages and can change; check them before you decide.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume