Video call avatar or scripted talking video: which do you need?
A real-time avatar answers people live; a scripted talking video is a file you render and review first. How to choose, and what Sume Avatar 1.0 covers.
Choose a scripted talking video when the words are known in advance and you want to review the result before anyone sees it. Choose a real-time video-call avatar only when the viewer must be able to talk back. Sume Avatar 1.0 covers the first case: it renders a talking video from a script through a job, and it does not list a live conversational model.
The two shapes of product
The difference is the loop. In a real-time system the model listens, decides when to speak, and generates frames as the call runs. Tavus describes its Griffin research preview that way: it generates every frame in real time from one reference image, and it can interrupt or be interrupted (Tavus: Griffin, read 2026-10-08). That preview is limited to trusted testers.
A scripted video is a batch job. You submit text, Sume returns a job, you poll it, and you download the finished MP4. If the take is wrong you change the script and run again.
| Question | Real-time call avatar | Scripted talking video |
|---|---|---|
| Viewer can talk back | Yes | No |
| You can review before publishing | No, it happens live | Yes, the output is a file |
| Typical use | Tutoring, practice calls, support | Ads, explainers, onboarding, product demos |
| On Sume | Not listed | Avatar 1.0 talking-video route |
What the scripted route accepts
From the Generate avatar video page (read 2026-10-08):
- One of script or video_inputs, never both.
- Estimated duration of 4-60 seconds inclusive; longer scripts must be shortened or split into jobs.
- Aspect ratios 1:1, 3:4, 9:16, 4:3 and 16:9, with 9:16 as the default.
- Resolution is 720p at this time.
- Optional product_image, scene direction and inline captions.
Where the risk sits
Live avatars shift risk to runtime: whatever the model says in the call is what the user hears. Scripted videos move the risk to the writing stage, where a person can read the script. That is why most commercial presenter work stays scripted.
Tavus itself notes the same properties that make such models natural interfaces also let them deceive a person into believing they are not talking to an AI, and says it is working on disclosure features. Whichever shape you pick, decide where you will disclose that the presenter is AI before you publish.
To start with the scripted route, follow Jobs and results: submit async, poll status_url until terminal is true, then read result_url.
A short decision list
Use a scripted video when the content is promotional, instructional or compliance-sensitive, since the script can be checked by legal or brand teams before rendering. Use a live avatar when the value is the conversation itself, such as practice calls or tutoring, and accept that you cannot pre-approve each reply.
Cost and latency also differ. A job-based render has a clear unit: you submit one request and get one file. A real-time session is a stream whose cost model depends on the vendor, and Tavus gives no price for Griffin-Lite on its page. Sume's avatar price per second by tier is published in its rate card, which makes budgeting a batch of scripted clips straightforward.
Finally, plan the handoff. Many teams use both: a scripted avatar video to explain a product, then a human or a live agent for questions. Nothing in Sume's docs prevents that split, and the scripted half is the part Sume ships.
Sources
Related posts
More in Sume Avatar 1.0
- Recorded voice to talking face: Fabric or H3 Max lip sync on Sume?
Sume has two still-plus-audio routes: veed/fabric-1.0 for 1-300 s and MiniMax H3 Max Lip Sync for audio of 5-14.8 s. Choose by audio length and what you have.
- Reuse one AI spokesperson across videos with an avatar handle
A Sume avatar_handle is a stable name for a ready avatar. Sume strips a leading @, so you create once and call the handle in every talking-video request.
- Add a silent beat to an AI avatar video with voice type silence
In Sume multi-scene avatar videos a scene with voice type silence is a pause with no speech. It needs a duration and rejects script text. Rules and example.
- Sume Avatar 1.0 ratios vs TikTok ad ratios: three overlap, two do not
Avatar 1.0 makes 1:1, 3:4, 9:16, 4:3 or 16:9 at 720p, 4 to 60 s. TikTok's non-Spark ad page lists 9:16, 16:9 and 1:1, so 3:4 and 4:3 need a crop.
Written by Sume