Make an AI avatar say exact words: Tavus echo mode vs Sume
Tavus echo mode sends text or audio straight to the avatar for playback, skipping perception and speech recognition. Sume's avatar video renders your script.
To make an avatar say exact words on Tavus, use echo mode: a simplified pipeline that accepts text or audio and plays it through the avatar, bypassing most conversation processing. On Sume, you send the words as the script of an avatar video and receive a finished MP4 after a job runs. Echo is for live delivery of text you control; Sume is for a file you can reuse.
Both avoid the usual live-agent risk: the model does not decide what to say.
Tavus pipeline modes
Tavus lists five ways to run a conversation: the full pipeline (the default, with all layers active), echo mode, a LiveKit agent integration, a Pipecat integration, and a custom LLM or your own logic. Echo mode is described as sending text or audio directly to the avatar for playback while skipping perception and speech recognition. The custom LLM mode is flagged as adding latency from external processing (Tavus docs: Pipeline Modes).
| Mode | What it does | Trade-off noted by Tavus |
|---|---|---|
| Full pipeline | All layers active; default | Low utterance-to-utterance latency with defaults |
| Echo | Text or audio played through the avatar | Bypasses perception and speech recognition |
| LiveKit agent | Tavus face in a LiveKit room | Not compatible with Tavus perception and speech recognition |
| Pipecat | Tavus joins as a transport or video service | Not compatible with the full Tavus multimodal stack |
| Custom LLM | Your endpoint writes the replies | Adds latency |
The Sume route to the same result
Sume's avatar video takes exactly one of script or video_inputs, with a ready avatar handle, a quality tier and an aspect ratio (default 9:16). Captions, if requested, are built from the spoken script text (Generate avatar video). The job is asynchronous: submit, poll GET /v1/jobs/{id}/status, and read the result when it completes, or use a webhook.
That is slower than playing text through a live face, but the output is a file. You can review it, caption it, trim it and reuse it without paying to speak the same words twice.
- Choose echo when the viewer is present and the text is generated at the moment, such as reading out an order status.
- Choose a rendered script when the same words go to many people or into an ad.
- Whichever you use, say that the speaker is synthetic where the law or the platform requires it.
Questions to ask
Does the viewer need the words now, or can they wait for a file? Is the text the same for many people? Will you need a caption, a trim or a re-render later? If the answers are later, many and yes, render once and reuse.
If the answer is now and unique, echo mode avoids waiting for a job. Check the latency on your own network, since the Tavus page names no figure for echo playback.
Writing a script that renders well
Write for the ear. Short sentences, numbers spelled the way you want them spoken, and one idea per scene. Read the script aloud once before you submit, since a clumsy sentence costs the same to render as a good one.
Keep the estimated duration inside the 4-60 second window and leave margin, because the estimate is what the service checks before it accepts the job.
Sources
Related posts
More in Comparisons
- Micro-drama lead: Avatar 1.0 or Seedance 2.5? Pick by dialogue
Pick Sume Avatar 1.0 for talking to camera and Seedance 2.5 for a lead who moves through places. Cost per minute: $14.70 against $34.67.
- Clipchamp free captions and silence removal vs a scripted cut list
Clipchamp's pricing page lists AI subtitles and silence removal in the free plan. When a scripted Sume cut list still makes sense, and what each step costs.
- Midjourney alternative for product stills by API: what to send on Sume
Sume's image catalog has no Midjourney model. For product stills it lists reference-based edit models instead; here is which id fits which job, and how to test.
- MiniMax H3 768p: $0.08/s on MiniMax's page, $0.06 in Sume's card
MiniMax's own page lists H3 768P at $0.08 a second; Sume's H3 rate card encodes a $0.06 list from an Aug 25 fal read, times 1.25. Compare a 5 second clip.
Written by Sume