One Sume avatar handle across talking video, TTS, stills and lip sync
A Sume avatar handle is accepted by talking-video, TTS voice selection, face-in-still images and the lip-sync routes. What each uses from it and what it costs.
Yes: once an avatar is ready, its handle is the identity you pass to the Avatar Video route, to Sume TTS as a voice selector, to Sume Image 1.0 for face-in-still ads, and to the VEED Fabric and MiniMax H3 Max lip-sync routes. Creating the avatar costs $0.95 once; every later use is billed by the route that uses it.
| Route | What it takes from the avatar | Rate |
|---|---|---|
| POST /v1/avatar-1.0/talking-video | The avatar as presenter | $0.184, $0.245 or $0.55 per second by tier |
| Sume TTS 1.0 | The avatar's voice, when voice.status is ready | Per the TTS rate card |
| Sume Image 1.0 | The identity still for face-in-still ads | Per the image rate card |
| veed/fabric-1.0 | Avatar still plus your audio | 480p $0.10, 720p $0.1875 per audio second |
| minimax/h3-max/lip-sync | Avatar still plus 5 to 14.8 s audio | 480p $0.0625, 768p $0.10, 1080p $0.20 |
Why use the handle and not the id
The docs tell you to use avatar_handle in generation requests. The durable avatar_id is read-only. Sume normalizes a leading @ and stores the handle without it, so @product_host and product_host are the same avatar. A handle gives your code a stable name that does not depend on a generated id.
The voice is part of the avatar
The TTS schema says that when you pass an avatar, Sume resolves that avatar's TTS voice at submit time, and that the selector works exactly when the avatar's voice.status is ready. List GET /v1/avatar-1.0/avatars and check the status before you build a pipeline on it. For a voice that feeds a lip-sync route, ask for wav output, since the TTS description recommends wav/pcm_s16le/44100 for avatar mux use.
So one avatar can give you a face, a voice and an identity for stills. That is a consistency benefit, not a streaming one: every route above produces a finished file.
Costs to plan
A presenter used in a 30 second talking video on standard costs 30 x 0.184 = $5.52. The same face in a 10 second lip-sync of your own audio at H3 Max 768p costs 10 x 0.10 = $1.00. Add the one-time $0.95 for the handle. Always send an Idempotency-Key, so a retry does not create a second charge.
Naming it well
Handles are short names with their own rules, so choose one you will not want to change. A name that describes the role, such as the host of a series, ages better than a campaign name. Handles with a leading @ are stored without it, and some prefixes are reserved by Sume, so check the avatar creation page for the current rules before you design a naming scheme.
Keep a small table in your own system that records what each handle is for and whether the person it depicts has given permission to use it. Use of a real person's likeness is your responsibility.
Sources
Related posts
More in Sume Avatar 1.0
- Preview an AI avatar video before paying for the full render
Sume avatar video previews make first-frame stills only. Approve them, then call generate-video; you can change the final quality tier without a new preview.
- Put the AI disclosure in scene one: Sume avatar video_inputs recipe
A copyable Sume Avatar 1.0 request with a 3-second disclosure scene and a 12-second message. The 15 seconds cost $2.76 at standard, $3.68 at plus.
- Video call avatar or scripted talking video: which do you need?
A real-time avatar answers people live; a scripted talking video is a file you render and review first. How to choose, and what Sume Avatar 1.0 covers.
- Recorded voice to talking face: Fabric or H3 Max lip sync on Sume?
Sume has two still-plus-audio routes: veed/fabric-1.0 for 1-300 s and MiniMax H3 Max Lip Sync for audio of 5-14.8 s. Choose by audio length and what you have.
Written by Sume