Tavus screen share: the agent sees your screen; Sume takes images
Tavus screen share needs raven-1 perception, a live video room and a user who starts sharing. A Sume avatar clip takes a product image or scene photo instead.

Tavus screen share lets a live agent look at what a user shares and answer questions about it, but only with perception_model set to raven-1, inside a live video conversation, after a participant starts the share. Sume has no live viewing. An avatar video takes images you supply in advance: an optional product_image, a scene photo, or a per-scene background image.
So the choice is whether the agent needs to see something that changes during the conversation or you know the visual in advance.
What Tavus requires
Per Tavus, screen perception requires Raven, a live video room (not chat-only or audio-only) and a participant who starts sharing through the browser prompt. Raven can read text, spot errors and answer prompted visual questions about the shared screen. For structured goals, you use objectives with a visual modality and visual_awareness_queries. Screen share is enabled by default, and you restrict it by leaving the share button out of your UI (Tavus docs: Screen Share).
| Need | Tavus | Sume avatar video |
|---|---|---|
| Agent sees a live screen | Screen share with raven-1 | Not available |
| Product shown beside the avatar | Shared screen or Canvas image | Optional product_image (public HTTPS) |
| Scene set from a photo | Not documented on the page | scene: { type: "photo", image_url } |
| Per-scene backdrop | Background customisation | video_inputs[].background with type: "image" and a url |
The Sume inputs
Sume's avatar video accepts product_image for a product or reference image, scene with a prompt or a photo for scene direction, and per-scene image backgrounds in video_inputs. All media fields must be fetchable public HTTPS URLs; localhost, private-network, non-HTTPS and signed or private URLs are not accepted (Media inputs).
For a software demo, export the screen states you need as stills, host them at public URLs and assign one per scene, or record the demo separately and join it to the avatar clip with Timeline.
- Check a finished clip with
video_inspect: it returns stills from the video so you can confirm the right image appears in the right scene. - Do not put a private dashboard in an image you publish; the image URLs must be public, and the finished clip can be shared by URL.
Practical preparation
For a rendered clip, prepare images at the aspect ratio of the final video, host them at stable public HTTPS URLs, and keep the originals. A product shot on a clean background works as a product_image; a lifestyle photo works as a scene.
For a live session, test the share prompt in each browser your users have, and decide in advance what the agent should say if the user declines to share.
What can go wrong
Sume's media inputs docs say signed or private URLs are not accepted, so test each URL from outside your network before you submit. A very busy screenshot may read poorly at phone size, so crop to the area that matters.
A live share can expose more than you meant to show, because the user chooses what to share. A prepared image cannot, which is one reason regulated teams prefer scripted clips.
Sources
Related posts
More in Comparisons
- Tavus Sparrow-2 turn-taking vs writing pauses in a Sume clip
Tavus lets you tune turn_taking_patience and pal_interruptibility on a live agent. A Sume avatar clip has no turns: you author pauses as silence scenes.
- Together AI speech-to-text at $0.0015 a minute vs Sume STT at $0.01
Together AI lists Whisper Large v3 at $0.0015 per audio minute. Sume STT is $0.01 per minute with a 10-minute cap. Cost of 1,000 minutes, and what the gap buys.
- Together AI TTS runs $4 to $65 per million characters; Sume is $47.50
Together AI lists text-to-speech from $4 to $65 per million characters. Sume TTS is $0.0475 per 1,000, or $47.50 per million. Cost of 100 scripts on each.
- Transcription cost: Scribe v2 $0.22/hour vs Sume's $0.01 per minute
ElevenLabs Scribe v2 lists $0.22 per hour of audio; Sume's video inspect transcript is $0.01 per audio minute ($0.60 per hour). Cost of 10, 60 and 600 minutes.
Written by Sume