UGC-style ad batch: ten 12-second hooks on Avatar 1.0 for $29.40
Ten 12-second UGC-style hook variants cost about $29.40 on Sume Avatar 1.0 plus, $22.08 on standard. The request, captions and checks.
Ten UGC-style 12-second hook variants cost about $29.40 on Sume Avatar 1.0 at the default plus tier, $22.08 on standard, and $66.00 on max, before any product-image premium. The avatar itself is a one-time $0.95.
The batch is cheap because UGC ads are short and share almost everything. One avatar, one scene, one structure, and ten different hook lines.
One structure, ten hooks
A UGC ad in the docs example is three beats in one video_inputs plan: a 3-second hook, a 4-second silent demo beat and a 5-second call to action, 12 seconds in total. The scene backgrounds describe a casual bedroom with native UGC lighting. The current execution supports one resolved avatar per final video and expects the backgrounds to resolve to one shared scene. So a batch varies the hook text, not the location.
Rates are the charged per-second rates from the Sume pricing package, by quality tier: standard $0.184, plus $0.245 and max $0.55 without a product.
| Tier | Per second | One 12 s variant | Ten variants |
|---|---|---|---|
| standard | $0.184 | $2.208 | $22.08 |
| plus (default) | $0.245 | $2.94 | $29.40 |
| max | $0.55 | $6.60 | $66.00 |
| plus, plus one-time avatar | $30.35 |
What to vary and what to hold
Keep the demo beat as voice.type: "silence", which the docs describe as a beat with no speech and a required duration. Spoken beats use type: "text" with one of script or input_text, not both. Turn on inline captions with a style such as slam or punch when the ad will play muted. Inline captions do not create a separate billed caption job, so they add no line to the batch cost, and a caption failure is soft: the avatar job can still succeed with a clean video.
The hook is the line to test. Write ten openers around one offer, for example a price drop, a gift deadline, a stock warning, a question, a reaction. Keep the demo and the call to action identical so that any difference in results comes from the hook.
One variant
Here is a hook-only variant of the multi-scene request. Only the first voice script changes between files.
curl -X POST https://api.sume.com/v1/avatar-1.0/talking-video \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: ugc-hook-03" \
-d '{
"avatar_handle": "product_host",
"aspect_ratio": "9:16",
"quality": "plus",
"captions": {"enabled": true, "style": "slam", "language": "auto"},
"video_inputs": [
{"id": "hook", "voice": {"type": "text", "script": "Your gift list just got shorter.", "duration": 3},
"background": {"type": "prompt", "prompt": "Casual bedroom, native UGC lighting"}},
{"id": "demo", "voice": {"type": "silence", "duration": 4},
"background": {"type": "prompt", "prompt": "Casual bedroom, native UGC lighting"}},
{"id": "cta", "voice": {"type": "text", "script": "Free shipping until Monday.", "duration": 5},
"background": {"type": "prompt", "prompt": "Casual bedroom, native UGC lighting"}}
]
}'Before you run ten
Check three things before you spend ten renders. Run one variant at standard and watch it, since the lip movement and the scene decide whether the avatar reads as native UGC. Make sure the planned seconds add up: voice durations of 3, 4 and 5 give 12, and Sume accepts plans that estimate between 4 and 60 seconds. And label synthetic presenters where the platform or the law requires it, since a UGC look is the whole point of the format and also the reason a label matters.
Treat the figures as estimates from the rate card. The job result is the billing record. Once the batch finishes, kill the bottom seven hooks and re-render the top three at max if the campaign deserves it: that adds $19.80 for three 12-second clips, $6.60 each.
How batches go wrong
A UGC batch fails in the same few ways, and most of them are about the scripts, not the render. Hooks that are too alike teach you nothing, so make each opener a different claim. Scripts that are too long do not fit the three-second slot: at normal speech pace three seconds holds a short sentence, so keep hooks under about eight words and set the duration to match.
Another trap is changing the demo beat between variants. The silent demo is the part that holds the product on screen. If you change it, you are testing two things at once and the result tells you nothing about the hook.
The last trap is skipping the review. Ten renders at plus is $29.40, and one bad avatar pose multiplied by ten is a wasted batch. Preview first if the scene is new. An approved first frame is reused by the final render, and the preview docs say a downgrade or upgrade of the tier needs no new preview.
Past ten variants
To scale beyond ten, wrap the single request in a loop with one Idempotency-Key per variant and poll each job. For a hundred variants, move the same recipe into a Format and use a bulk queue, which accepts up to 100 items. The posts linked below cover that route.
Sources
Related posts
More in Sume Avatar 1.0
- UGC-style avatar ad: hook, silent demo beat and CTA in video_inputs
Build a three-scene UGC avatar ad with POST /v1/avatar-1.0/talking-video: a spoken hook, a silent demo beat, a spoken CTA, inside the 4 to 60 second window.
- What Sume Avatar 1.0 Does and Does Not Do vs Live Avatars
A plain list of what Sume Avatar 1.0 renders (scripted 4-60 s clips) and what it does not do (real-time conversation), set beside Tavus Griffin's live model.
- Can I use Tavus Griffin-Lite yet? Preview status and what to ship
Tavus Griffin-Lite is a research preview for select trusted testers, not open to customers. Here is what the page says, and what you can build on Sume today.
- Introducing Sume Avatar 1.0
Sume Avatar 1.0 is a multi-agent orchestration system as a single avatar model.
Written by Sume