Create an AI spokesperson avatar from a prompt or photo
Sume Avatar 1.0 creates a reusable avatar at a flat $0.95 from a text prompt, a photo, or props. How to make one before you generate talking videos.
Create the spokesperson once with POST /v1/avatar-1.0/generate, then reuse its avatar_handle for every ad. Avatar creation is a flat $0.95 per avatar (list prices from Sume's catalog rate card (GET /v1/catalog), confirmed in the repo's catalog file read 2026-10-03). The input is a prompt, a photo, or props, per Avatar.
Input types
The Jobs docs show the prompt form as {"avatar_handle":"studio_presenter","input":{"type":"prompt","prompt":"Friendly studio presenter"}}. The Avatar doc covers the props type as well.
| Input type | Use when |
|---|---|
| prompt | You want a new presenter from a description |
| photo | You have a source face image |
| props | You want items in the presenter's hands or frame |
Retry safely
Send an Idempotency-Key on the submit. Sume's jobs docs say to reuse the same key only for the same operation and payload, so a network retry does not create a second avatar. See Jobs and results.
Then generate
Once the avatar is ready, generate avatar video takes avatar_handle plus a script or video_inputs. Billing for those videos is per second, separate from the $0.95 creation fee.
Sources
Related posts
More in Sume Avatar 1.0
- Demand Gen: what to hold constant between two AI video ads
Google says Demand Gen experiment campaigns should differ in one variable. A field-by-field checklist for two Sume avatar videos so only the hook changes.
- Does Sume face swap keep the original voice, outfit and scene?
Sume's Face Swap (Beta) keeps the source clip's audio, camera and timing but replaces the whole person, not just the face. What stays, what changes, cost.
- Face swap for UGC ads: source clip rules (Sume beta)
Sume's beta Avatar Face Swap applies a ready avatar face to a public 4-15 second source video with audio. Required fields, quality tiers and what it rejects.
- How to make an AI UGC ad look less staged with Avatar 1.0
Less-staged AI UGC comes from the first frame: phone-style framing, a casual scene prompt, an approved preview, then the final render. The levers Sume exposes.
Written by Sume