Make an AI avatar say exact words: Tavus echo mode vs Sume

Tavus echo mode sends text or audio straight to the avatar for playback, skipping perception and speech recognition. Sume's avatar video renders your script.

5 min readSume
All posts

To make an avatar say exact words on Tavus, use echo mode: a simplified pipeline that accepts text or audio and plays it through the avatar, bypassing most conversation processing. On Sume, you send the words as the script of an avatar video and receive a finished MP4 after a job runs. Echo is for live delivery of text you control; Sume is for a file you can reuse.

Both avoid the usual live-agent risk: the model does not decide what to say.

Tavus pipeline modes

Tavus lists five ways to run a conversation: the full pipeline (the default, with all layers active), echo mode, a LiveKit agent integration, a Pipecat integration, and a custom LLM or your own logic. Echo mode is described as sending text or audio directly to the avatar for playback while skipping perception and speech recognition. The custom LLM mode is flagged as adding latency from external processing (Tavus docs: Pipeline Modes).

Tavus CVI pipeline modes (read 2026-10-03)
ModeWhat it doesTrade-off noted by Tavus
Full pipelineAll layers active; defaultLow utterance-to-utterance latency with defaults
EchoText or audio played through the avatarBypasses perception and speech recognition
LiveKit agentTavus face in a LiveKit roomNot compatible with Tavus perception and speech recognition
PipecatTavus joins as a transport or video serviceNot compatible with the full Tavus multimodal stack
Custom LLMYour endpoint writes the repliesAdds latency

The Sume route to the same result

Sume's avatar video takes exactly one of script or video_inputs, with a ready avatar handle, a quality tier and an aspect ratio (default 9:16). Captions, if requested, are built from the spoken script text (Generate avatar video). The job is asynchronous: submit, poll GET /v1/jobs/{id}/status, and read the result when it completes, or use a webhook.

That is slower than playing text through a live face, but the output is a file. You can review it, caption it, trim it and reuse it without paying to speak the same words twice.

  • Choose echo when the viewer is present and the text is generated at the moment, such as reading out an order status.
  • Choose a rendered script when the same words go to many people or into an ad.
  • Whichever you use, say that the speaker is synthetic where the law or the platform requires it.

Questions to ask

Does the viewer need the words now, or can they wait for a file? Is the text the same for many people? Will you need a caption, a trim or a re-render later? If the answers are later, many and yes, render once and reuse.

If the answer is now and unique, echo mode avoids waiting for a job. Check the latency on your own network, since the Tavus page names no figure for echo playback.

Writing a script that renders well

Write for the ear. Short sentences, numbers spelled the way you want them spoken, and one idea per scene. Read the script aloud once before you submit, since a clumsy sentence costs the same to render as a good one.

Keep the estimated duration inside the 4-60 second window and leave margin, because the estimate is what the service checks before it accepts the job.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume