Tavus Video to Face replica vs a Sume avatar from a photo or prompt
Tavus builds a replica from video or a photo; Sume builds an avatar from a prompt, props or a public photo URL. What each input gives you, and what it costs.
The question: how do I get a digital person from what I already have?
Tavus documents two routes to a custom replica: Video to Face and Image to Face, through PAL Maker or the API (Tavus replica docs, read 2026-10-05). Training from video captures movement and voice. A photo gives a face, and Tavus says the voice must come separately.
Sume Avatar 1.0 has three input types on one endpoint: a text prompt, structured props (ethnicity, sex, age) or a photo. The photo must be a public HTTPS image URL. The result is an avatar you name with an avatar_handle and reuse in talking videos.
| Question | Tavus replica | Sume avatar |
|---|---|---|
| Starts from a video of a person | Yes (Video to Face) | No, no video input type |
| Starts from one photo | Yes (Image to Face) | Yes (type photo, public HTTPS image_url) |
| Starts from no source at all | Stock replicas (100+ listed) | Yes (type prompt or props) |
| Voice | Captured from video; photo needs a separate voice | Avatar voice must be ready before a video |
What Sume charges for the avatar itself
Creating a Sume avatar is a flat $0.95. A video then bills per second of the finished clip: $0.184, $0.245 or $0.55 for standard, plus and max (no product image). So a 10 s plus clip is $2.45, and the avatar plus that first clip is $3.40.
Tavus gates custom faces by plan: Starter, Growth and Enterprise per its docs. We did not find a per-replica price on the pages we read, so compare plan prices, not a per-face figure.
Pick by what you hold today
- You have footage of a presenter and want their motion and voice captured: Tavus Video to Face fits that input.
- You have one clean photo or none: Sume takes the photo, props or a prompt and returns a handle.
- You need a fixed spoken script as a clip, not a live conversation: use the Sume talking-video route with the handle.
Rights come first
Tavus states you are responsible for having the necessary rights and permissions for the footage you train on. The same applies to any photo you send to Sume: use only images of people who agreed.
Sources
Related posts
More in Sume Avatar 1.0
- What an AI avatar may not claim on TikTok Shop
TikTok Shop prohibits AI that impersonates real people or invents doctors and experts to endorse products. An avatar can present, but it can't be a fake expert.
- UGC-style ad batch: ten 12-second hooks on Avatar 1.0 for $29.40
Ten 12-second UGC-style hook variants cost about $29.40 on Sume Avatar 1.0 plus, $22.08 on standard. The request, captions and checks.
- UGC-style avatar ad: hook, silent demo beat and CTA in video_inputs
Build a three-scene UGC avatar ad with POST /v1/avatar-1.0/talking-video: a spoken hook, a silent demo beat, a spoken CTA, inside the 4 to 60 second window.
- What Sume Avatar 1.0 Does and Does Not Do vs Live Avatars
A plain list of what Sume Avatar 1.0 renders (scripted 4-60 s clips) and what it does not do (real-time conversation), set beside Tavus Griffin's live model.
Written by Sume