Tavus Magic Canvas cards vs what a rendered Sume avatar clip shows
Tavus Magic Canvas shows 8 kinds of interactive cards in live video calls only. A Sume avatar clip is a fixed MP4: CTA goes in a closing scene.
Tavus Magic Canvas lets a live agent show interactive cards, such as multiple-choice questions, calendars and charts, beside the video. It works only in video conversations. A Sume avatar clip is a finished MP4 and cannot take clicks; the closest equivalent is a closing call-to-action scene in the video and an interactive element on the page around it.
That split is the practical answer: what must respond to the viewer lives on the page, and what must stay the same lives in the file.
What Magic Canvas offers
Tavus lists eight components. Interactive: question, input, calendar and scheduling_embed. Display-only: text, image, chart and alert. The PAL decides when to show a card, responses go back to the PAL and your webhook, and cards render in a sandboxed iframe on a side rail. Audio-only, text-chat and external-meeting conversations do not get Canvas actions (Tavus docs: Magic Canvas).
| Tavus component | Kind | In a Sume avatar clip |
|---|---|---|
| question, input | Interactive | Ask on the page next to the video |
| calendar, scheduling_embed | Interactive | Link or embed on the page |
| text, alert | Display | Closing scene text or burned captions |
| image | Display | product_image, a scene photo or an image background |
| chart | Display | Show as a still in a scene, or on the page |
Building the closing scene
A multi-scene avatar video takes ordered video_inputs, each with its own voice and optionally a background, and the docs include a final scene with id cta. Current execution supports one resolved avatar per final video and expects scene backgrounds to resolve to one shared scene, so plan the visual as one setting with changing speech rather than different locations (Generate avatar video).
If you need pieces from different sources joined, Timeline 1.0 assembles ordered video slots over one audio track (Timeline 1.0).
- Say the action out loud in the last scene and repeat it on screen.
- Put the link, booking widget and form in the page, not in the video.
- Caption the clip so the call to action reads with sound off.
When a live card is worth it
If the next step depends on an answer the viewer gives, such as picking a plan or a time slot, a live card shortens the path. If the next step is the same for everyone, a clip and a button do the job at lower complexity.
Magic Canvas needs a video conversation, so check that your channel supports one. For email, ads and social feeds, a rendered clip with a clear closing scene is the format that works.
Mixing the two
Some funnels use both. A rendered clip in an email gets the click, and the landing page hosts a live conversation with cards for people who want to go deeper. Keep the message the same in both so the handoff does not feel like a different company.
Measure each step separately: clip views, clicks to the page, and conversations started, so you know where people drop off.
Sources
Related posts
More in Comparisons
- Tavus max_call_duration is plan-capped; Sume's cap is 60 s a job
Tavus ends a call at max_call_duration, capped by your plan; an unjoined call times out after 300 s. Sume's avatar video takes 4-60 s per job. Table inside.
- Tavus max_participants counts the face: seats vs a shareable clip
In a Tavus conversation the AI face takes a seat: max_participants of 2 means one human plus one face. A Sume clip has no seats.
- Tavus Sparrow-2 turn-taking vs writing pauses in a Sume clip
Tavus lets you tune turn_taking_patience and pal_interruptibility on a live agent. A Sume avatar clip has no turns: you author pauses as silence scenes.
- Together AI speech-to-text at $0.0015 a minute vs Sume STT at $0.01
Together AI lists Whisper Large v3 at $0.0015 per audio minute. Sume STT is $0.01 per minute with a 10-minute cap. Cost of 1,000 minutes, and what the gap buys.
Written by Sume