Captions AI avatar looks vs one Sume avatar handle per look
Captions added Avatar Looks to save new looks per avatar. On Sume you create one avatar handle per look from a prompt, a profile or a reference image.
Captions' August 2026 notes say you can create and save new looks for avatars and use them across Prompt to Video and other avatar projects. Sume has no separate look object in the docs I read: to get a new look you create a new avatar, from a prompt, a profile or a reference image, and keep its handle.
Captions text is from its what's-new page and Sume's from the docs, both read 2026-10-01.
What are the three ways to make an avatar on Sume?
The Create new avatar page lists three inputs. Each request creates a job; poll it until it completes, then use the returned avatar handle or resource id to generate avatar videos.
| Input | What you provide |
|---|---|
| Prompt | Describe the avatar you want |
| Profile | Structured traits (the props input type in the API) |
| Image | A reference image (input.type: "photo") |
How do I keep several looks for one character?
Create one avatar per look and name the handles so the relationship is obvious, for example one handle for the casual look and one for the studio look. The route is POST /v1/avatar-1.0/generate with a top-level avatar_handle; a leading @ is allowed and Sume stores it without it. This is a naming convention on your side, not a documented look-grouping feature.
How do I use a look in a video?
Pass the handle as avatar_handle to POST /v1/avatar-1.0/talking-video. A product image is optional: omit product_image for a productless avatar video. See Generate avatar video. For comparison with another vendor's look packs, read HeyGen look packs vs a new Sume avatar per look.
Does Sume copy a person's likeness from a video upload?
The avatar docs describe prompt, profile and image inputs only. Check the page for current inputs before building a flow that depends on anything else.
Sources
Related posts
More in Use cases
- Multilingual video captions: language is a hint, not a font
On Sume video captions, `language` only tells speech-to-text what to expect. Look and font come from `style`, `design` and `font`, never from the language.
- Coupang main image: white background, square, 95% fill
Coupang's main image should be square on white with the item near 95% of the frame. Request aspect_ratio 1:1 from Sume and check the fill yourself.
- ElevenLabs agent hold audio: 180 s, 40 MB; make it with Sume
ElevenLabs agent hold audio must be MP3 or WAV, up to 40 MB and 180 seconds. Steer a Sume track's length in the prompt, trim with Timeline audio, export mp3.
- How long must the EU AI label stay on screen? Cue timing
The Commission gives no seconds: a limited-time disclosure should stay visible long enough to be read by people with cognitive or processing difficulties.
Written by Sume