Referral program launch video: avatar above the share screen
Explain a refer-a-friend offer in 30 seconds: an avatar script with a silent beat, the share screen stacked below it, and a plan check before you pay.
To explain a refer-a-friend offer in one short video, write a 4 to 60 second avatar script in three scenes (the offer, a silent beat, the ask), render it as a 9:16 Sume avatar video with inline captions, and stack a still of your share screen above the avatar with one Timeline compose job. The viewer hears the offer and sees where the code lives in the same frame.
A referral offer fails when people cannot picture the step. Saying "share your link from the account page" is weaker than showing the account page while someone says it. The avatar supplies the voice, and the still supplies the proof.
Scene plan
Sume's avatar video takes ordered video_inputs. Each scene has a voice of type: "text" (with script or input_text) or type: "silence", and a required duration. The total planned length has to land between 4 and 60 seconds. A silent scene is a beat with no speech, which is where a viewer's eye goes to the screen below.
Write the offer once, in plain words, with the real reward and the real condition. Put numbers you can defend in the script, because the spoken words and the caption text come from the same source.
Request
The body below follows the multi-scene example in the avatar video docs, with a referral script. Replace the avatar handle with one that is ready in your workspace and keep an Idempotency-Key on every submit. Aspect ratio defaults to 9:16, resolution is 720p, and quality is plus unless you send standard or max.
curl -X POST https://api.sume.com/v1/avatar-1.0/talking-video \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: referral-launch-001" \
-d '{
"avatar_handle": "YOUR_AVATAR_HANDLE",
"aspect_ratio": "9:16",
"captions": { "enabled": true, "style": "slam", "language": "auto" },
"video_inputs": [
{ "id": "offer", "voice": { "type": "text", "script": "Give a friend twenty dollars, get twenty dollars.", "duration": 5 },
"background": { "type": "prompt", "prompt": "Bright kitchen, natural light" } },
{ "id": "look", "voice": { "type": "silence", "duration": 4 },
"background": { "type": "prompt", "prompt": "Bright kitchen, natural light" } },
{ "id": "ask", "voice": { "type": "text", "script": "Open Account, tap Share, send your link.", "duration": 5 },
"background": { "type": "prompt", "prompt": "Bright kitchen, natural light" } }
]
}'Stack the share screen
When the avatar job is done, GET /v1/jobs/:id/result returns the video on media.sume.com. Import your share-screen still with POST /v1/media-imports, because Timeline compose accepts only this workspace's hosted files. Then submit a stack job. The still takes the share of the frame that you set with ratio (0.1 to 0.9), and the video takes the rest. The job is $0.02 flat per the docs (read 2026-10-10), and its output length always comes from the video layer.
Use output.width 720 and output.height 1280 to match the avatar clip. Compose sets the still on the top and the video below by default, which is the half-banner layout.
Check before you pay for max
Avatar video previews return first-frame stills for each scene, so you can approve the framing before the full render, and they never burn captions into those stills. Captions are burned into the avatar clip before the stack, so look at a frame of the stacked result before you publish: the captions must not cover the share screen's button or the code field. Run POST /v1/avatar-video-previews, then generate-video on the preview id once the stills look right, per the preview docs.
What each stage returns
Poll each job with the shared envelope in Jobs and results, and store the job id before you do anything else.
| Stage | Endpoint | Constraint from the docs |
|---|---|---|
| Avatar scenes | POST /v1/avatar-1.0/talking-video | 4 to 60 s planned length, 720p |
| Inline captions | captions on the same request | Not a separate billed job; a caption failure is soft |
| First-frame check | POST /v1/avatar-video-previews | Stills only, no burned captions |
| Stack share screen | POST /v1/timeline-1.0/compose | $0.02 flat, hosted files only |
Keep the offer honest
Say the condition in the clip, not only in the fine print: who qualifies, when the reward arrives, and any limit. If the offer changes next month, re-render the stack job with a new still and keep the avatar clip, since the stack job is a flat two cents and needs no new avatar render.
One caution on codes. Do not ask the avatar to read a long code aloud. Show the path to the code on the still and let the viewer's own account supply the value.
Sources
Related posts
More in Use cases
- Daily special video for 21 cents: TTS, one still, one render
A restaurant daily special as a 12-second video on Sume: one still, 180 characters of TTS and a Timeline render. Cost per day and per month, read 2026-10-10.
- Detach 13 seconds of speech from a clip and lip-sync it on H3 Max
Sume can detach audio from a hosted clip by time range, then drive a talking face with H3 Max lip-sync. $1.31 for 13 seconds; rights and limits explained.
- RV Dealer Listing Video: Three Clips, Music and Captions for $2.32
An RV listing video from three 5-second Wan 3.0 clips at 720p, a music bed, a timeline render and captions costs $2.32 on Sume. Here is the build.
- School bake sale video: three small stills, free music, 53 cents
A bake sale video on Sume with three 0.5K stills, a free BGM track and date cues. Silent-spine Timeline render and caption cost, read 2026-10-10.
Written by Sume