Captions Avatar Looks in Prompt to Video: the Sume handle workflow
Captions saves looks inside one avatar for Prompt to Video. For several looks in Sume, make one avatar_handle per look and render one job per final video.
To reuse several looks across Sume videos, create one avatar per look, name the handles so you can tell them apart, and send the right avatar_handle with each video request. Captions, per its release notes, instead saves looks under an avatar for use across Prompt to Video and other avatar projects.
Captions facts are from its release notes; Sume facts from Create new avatar and Generate avatar video, read 2026-10-01. For the plain comparison see captions avatar looks vs a new avatar handle.
How do I organise several looks in Sume?
Pick a handle per look, for example host_studio and host_street. The handle may include a leading @; Sume stores it normalized without @, so do not rely on the @ to distinguish two handles. Create each with POST /v1/avatar-1.0/generate and a distinct Idempotency-Key, poll the job, then keep a small table of handle to purpose in your own app.
| Your label | avatar_handle | Created from |
|---|---|---|
| Studio | host_studio | Prompt |
| Street | host_street | Reference image (photo) |
| Product demo | host_demo | Profile (props) |
Which input kind fits a look?
Prompt describes the avatar from text, Profile gives structured traits, Image uses a reference image. The image_url must be a fetchable public HTTPS image URL, or the request is rejected before generation.
How do I pick a look per video?
Set avatar_handle on POST /v1/avatar-1.0/talking-video. Current execution supports one resolved avatar per final video, so a video cannot cut between two handles; render separate jobs and join them downstream.
Sources
Related posts
More in Use cases
- Captions horizontal 16:9 AI Edit vs Sume avatar aspect ratios
Captions AI Edit now outputs 16:9. Sume avatar videos accept 16:9 too, at 720p only, with a script or plan estimated at 4-60 seconds. 9:16 is the default.
- Captions app SRT export vs Sume burned-in caption cues
The Captions app can export an SRT file for external caption tracks. Sume burns captions into the MP4 from authored cues with text, start and end.
- Multilingual video captions: language is a hint, not a font
On Sume video captions, `language` only tells speech-to-text what to expect. Look and font come from `style`, `design` and `font`, never from the language.
- Chrome Web Store 440x280 promo tile: generate at 1760x1120
The Chrome Web Store small promo tile is 440x280, below GPT Image 2.5's pixel floor. Request 1760x1120 (exact 4x, same 11:7 ratio) on Sume and downscale.
Written by Sume