Live AI avatar for streaming: pre-rendered clips instead

Streamers asking for a live AI avatar get conversation APIs. Sume cannot go live, but renders avatar intros, stingers and ad reads. Tavus greenscreen vs Sume.

5 min readSume
All posts

Sume cannot give you a live AI avatar for a stream. It does not run a live session or return a video feed. What it can do is render short avatar clips of 4-60 seconds that you play from your streaming software: a channel intro, a sponsor read, a starting-soon loop, a scene that explains the rules of a giveaway. If you need an avatar that answers chat in real time, you need a live avatar API.

Most streamers who search for this want one of those two things, and they are different products.

What does the live route involve?

A live avatar vendor creates a conversation and gives you a feed or a room. Tavus's create-conversation reference lists a property apply_greenscreen, where the background "will be replaced with a greenscreen", which would let a streamer key the avatar over a scene (read 2026-10-03). The same reference sets max_call_duration to 3,600 seconds by default, so a long stream has to restart sessions. Tavus's Griffin is gated to trusted testers and not available to customers (Tavus Griffin post, read 2026-10-03).

What does Sume render for a stream?

A clip, with the avatar speaking your script over a scene you describe. aspect_ratio supports 1:1, 3:4, 9:16, 4:3 and 16:9, with 16:9 for a normal stream scene. Resolution is 720p. scene can be a text prompt or a photo reference, and inline captions burn the spoken words into the MP4 (Generate avatar video).

Sume's docs list no greenscreen or transparent-background option for avatar video, so the avatar arrives inside its generated scene rather than as a cut-out layer. If you need a cut-out, plan that in your own compositing.

Live avatar feed vs rendered clip for streams, read 2026-10-03
Stream needLive avatar (Tavus)Rendered clip (Sume)
Answers chat liveYesNo
Greenscreen backgroundapply_greenscreen propertyNot offered in the docs
Max continuous runmax_call_duration, default 3,600 sOne clip, 4-60 s
Reviewed before airNoYes, preview stills first
Repeats identicallyNoYes

How do I build a small clip library?

Write each clip as a script, preview it, then render the keepers. Reuse the same avatar_handle so the presenter is consistent, and put captions on for muted viewers. Every submit is a job with an Idempotency-Key, so your own scripts can re-run without duplicating renders.

curl -X POST https://api.sume.com/v1/avatar-video-previews \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: stream-intro-preview-001" \
  -d '{
    "avatar_handle": "sume_clawra",
    "script": "Welcome to the stream. Tonight we are playing the new season, and chat picks the next map.",
    "aspect_ratio": "16:9",
    "quality": "plus",
    "captions": {"enabled": true, "style": "slam", "language": "en"}
  }'

How long should a stream clip be?

Short. The 60-second ceiling is on Sume's duration estimate for the script, so a sponsor read that runs longer should be split into two jobs, or into scenes inside one job that stays within the window. Keep starting-soon loops and stingers short, which also keeps each render cheap to redo if a sponsor changes a line.

Because a clip is a plain MP4 with a public media.sume.com URL on completion, adding it to a scene in your streaming software is the same as adding any other video file. That step happens outside Sume; the docs do not describe a plugin for any streaming tool.

What if I need chat-reactive content?

Split it. Pre-render the parts that never change, such as intros and rules, and handle chat-reactive moments with a person or a live avatar service. Sume can help between those moments with captioned recaps: render a short avatar clip after the stream that summarises the highlights, written from your own notes, and post it as a short.

If you are weighing a live avatar vendor, look at the limits first. Concurrency is one stream on Tavus's free plan and up to 10 on Growth, and sessions default to 3,600 seconds (Tavus pricing and reference, read 2026-10-03). Those numbers matter more for a stream than a vendor's headline latency.

What should I disclose to viewers?

If a stream shows an AI-generated presenter, say so on screen or in the clip's first line. Rendered clips are reviewed in advance, which makes it easy to include the line. A live AI avatar answering chat is harder to control, and Tavus cites that deception risk as a reason for holding Griffin back.

  • Intro, outro and sponsor reads: render once, reuse every stream.
  • Answering chat live: needs a live avatar API, not Sume.
  • Muted autoplay on clips: turn on inline captions.
  • Keep each clip under the 60-second ceiling or split it into scenes.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume