Tavus 50 new Phoenix-4.5 stock faces vs a Sume avatar handle
Tavus added 50 first-party Phoenix-4.5 stock faces, 28 of them Pro. On Sume you create your own avatar from a prompt, profile or image and reuse its handle.
Tavus's changelog lists 50 additional first-party stock faces built on Phoenix-4.5, including 28 Pro faces. The Sume docs reserve the sume_ handle prefix for Sume system and default avatars (their example uses sume_clawra) but I found no catalog list of stock faces; you can create a reusable avatar from a prompt, a structured profile or a reference image, then call it by avatar_handle.
The Tavus entry is undated in its changelog snapshot, read 2026-10-01. Sume facts are from Create new avatar and Generate avatar video.
What did Tavus add?
The entry reads: "50 new Phoenix-4.5 stock faces: 50 additional first-party faces built on Phoenix-4.5, including 28 Pro faces." The changelog gives no price, availability date or API field for them, so this post does not either.
How do you get an avatar on Sume?
Avatar creation has three input kinds, sent to POST /v1/avatar-1.0/generate with a top-level avatar_handle and an input union.
| Input | What you provide |
|---|---|
| Prompt | Text describing the avatar you want |
Profile (props) | Structured traits such as ethnicity, sex and age |
| Image | A reference image at a public HTTPS URL |
How do I use the avatar afterwards?
Each request creates a job; poll it, then use the returned handle in POST /v1/avatar-1.0/talking-video. The handle may include a leading @, and Sume stores it normalized without the @. Current execution supports one resolved avatar per final video.
Stock face or created avatar: how do I decide?
A stock face saves the step of making one. A created avatar lets you describe the person you need and keep it as a handle across videos. If you have a photo of the person, the image route is covered in Tavus photo-to-replica versus a Sume photo avatar.
Sources
Related posts
More in Models
- VEED lipsync-v2 video + audio vs Sume still + audio lip sync
fal lists veed/lipsync-v2 as video plus audio in. Sume lip sync routes start from a still and an audio clip, so existing footage is not re-synced.
- Veo 3.1 only makes 16:9 and 9:16; which Sume video models add more
Google's Veo 3.1 supports 16:9 and 9:16. On Sume, Seedance, MiniMax and Wan add 4:3, 1:1 and 3:4, Kling adds 1:1, and Grok takes no ratio.
- Veo 3.1 prompts cap at 1,024 tokens; Sume's Omni at 20,000 characters
Google caps a Veo 3.1 text prompt at 1,024 tokens. On Sume, gemini-omni-flash-1.1 documents a 20,000-character cap. Tokens and characters differ.
- Vidu Q2 Pro Fast image-to-video 1080p vs Sume first frame
QwenCloud lists vidu/viduq2-pro-fast_img2video at 720P and 1080P. Sume has no Vidu id; send image_url as the first frame to a 1080p model.
Written by Sume