UGC-style ad on Sume: a Format run or the avatar endpoint
Two ways to make a UGC-style ad on Sume: a catalog Format run for a full cut, or the avatar talking-video endpoint for a 4 to 60 second presenter clip.
For a UGC-style ad, Sume gives you two calls. The avatar endpoint, POST /v1/avatar-1.0/talking-video, renders one presenter saying your script in 4 to 60 seconds. A Format run, such as the catalog's sume-close-camera-ugc, hands a brief to an agent that plans and assembles the cut with the generation tools. Pick the avatar endpoint when you control the words and want a predictable clip; pick a Format when you want the whole ad assembled from a brief.
Sources: Sume's Generate avatar video, Format catalog and Calling a Format.
How the two differ
The avatar route is a model call with a fixed shape. A Format run is one unattended agent turn in a fresh sandbox, ending in a receipt with media and, optionally, JSON in your schema. The Format docs say long-form host video typically finishes in 15 to 30 minutes, so design for async.
| Question | Avatar talking-video | Format run (catalog or your own) |
|---|---|---|
| Input | Avatar handle plus script or video_inputs | instruction plus free-form input |
| Who plans the shots | You, through scenes | The Format's recipe and the agent |
| Length | 4 to 60 seconds per job | Set by the recipe |
| Preview step | Yes, first-frame previews | No equivalent |
| Spend control | Quality tier, preview first | generation_spend_cap_usd per run |
| Batches | Submit jobs yourself | Bulk runs: 1 to 100 items, concurrency 1 to 16 |
Read the Format before you call it
The catalog says to read a Format with GET /v1/formats/sume/{slug} first, which returns its description and the io profile that declares the input and output kinds. Both can be null for older Formats, which means not declared. Do not guess the input keys.
curl -sS "https://api.sume.com/v1/formats/sume/sume-close-camera-ugc" \
-H "Authorization: Bearer $SUME_API_KEY"A practical split
Many teams use both. The Format makes a first cut of a new concept; once a hook works, the avatar endpoint repeats it with different scripts and keeps the face constant. Whichever you use, label synthetic presenters honestly and never present a generated person as a real customer.
- New concept, unknown shots: Format run.
- Known script, many variants: avatar endpoint with
video_inputs. - Many recipients: bulk runs with one idempotency key per batch.
Bottom line
Choose on who decides the shots. The more you want to control, the more the avatar endpoint fits; the more you want delegated, the more a Format fits.
Sources
Related posts
More in Use cases
- Vacation rental check-in video: an avatar host with silence beats
One avatar, one shared background and silence beats let a host walk guests through check-in in 60 seconds or less. Set video_inputs on Avatar Video.
- Vendor pages to read before AI music goes in an ad
The vendor pages that answer commercial-use questions for Suno v6, ElevenLabs Music v2.5 and Google Lyria 3.5, and which question each answers, read 2026-10-04.
- Vietnam AI Law Article 11: labelling real people, real events
Under Vietnam's AI Law, deployers must label AI media that imitates a real person's face or voice or recreates a real event. Art and film get a softer rule.
- Vietnam AI Law Article 11: machine-readable marks for AI media
Vietnam's Law No. 134/2025/QH15, in force 1 March 2026, makes AI providers mark audio, image and video in machine-readable form. Sume docs show no such mark.
Written by Sume