Personal trainer class promo: avatar hook, silent demo beat, CTA
Build an 11-second class promo for a fitness studio from three ordered scenes on one avatar, with a silence beat for the demo. Limits and request body included.
A personal trainer or small studio can make a class promo as three ordered scenes on one avatar: a spoken hook, a silent beat for the demo, and a spoken call to action. Sume's avatar talking-video route takes ordered video_inputs for this, with a planned total of 4 to 60 seconds.
The silent beat is the useful trick. A voice.type of silence with a duration gives you a gap with no speech where you can lay your own gym footage later, without the avatar talking over it.
Why scenes beat one long script
This month's trend roundups say avatars are mainly used for training, support and localized sales material, and that agents are being used to draft scripts and variants (AI Video Generation Trends, read 2026-10-07). A class promo is a small version of the same job: a repeatable structure with new words each week.
A single script gives you one continuous read. video_inputs lets you set the length of each beat, so the hook stays under four seconds and the call to action gets the room it needs.
The request
Create the avatar once with POST /v1/avatar-1.0/generate, then reuse its handle. In the talking-video body, send either script or video_inputs, never both. Spoken scenes use voice.type: "text" with one of script or input_text. A silent scene needs a duration and no text.
The current execution supports one resolved avatar per final video, and expects the scene backgrounds to resolve to one shared scene, so keep the background prompt the same across scenes.
curl -X POST https://api.sume.com/v1/avatar-1.0/talking-video \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: spin-class-promo-001" \
-d '{
"avatar_handle": "studio_coach",
"aspect_ratio": "9:16",
"quality": "plus",
"video_inputs": [
{ "id": "hook", "voice": { "type": "text", "script": "Thursday 6 a.m. spin is back.", "duration": 3 },
"background": { "type": "prompt", "prompt": "Bright studio, mirrors, soft morning light" } },
{ "id": "demo", "voice": { "type": "silence", "duration": 4 },
"background": { "type": "prompt", "prompt": "Bright studio, mirrors, soft morning light" } },
{ "id": "cta", "voice": { "type": "text", "script": "Book your first class free from the link below.", "duration": 4 },
"background": { "type": "prompt", "prompt": "Bright studio, mirrors, soft morning light" } }
]
}'Pick the quality tier
Quality is standard, plus or max. plus is the default and the balanced path, standard is the fastest Sume path, and max is the highest tier with slower turnaround. For a weekly promo that a trainer will post and replace, standard or plus is usually the practical choice.
Poll the job at /v1/jobs/{id}/status and fetch /v1/jobs/{id}/result when it is ready.
| Rule | Value | Why it matters for a class promo |
|---|---|---|
| Total planned length | 4 to 60 seconds | Hook + demo + CTA fits easily |
| script vs video_inputs | Send one, not both | Use video_inputs for beats |
| Silence scene | voice.type silence plus duration | Room for your own footage |
| Avatars per video | One resolved avatar | Keep one coach |
| quality | standard, plus (default), max | plus is the balanced default |
Where to add the real footage
Sume does not film your class. To put real class footage under the silent beat, export the avatar clip and edit it in your own editor, or assemble clips with Sume's timeline render. Disclose that the presenter is an AI avatar where your market or the platform expects it.
Sources
Related posts
More in Use cases
- Photo to talking selfie clip: put the spoken line in the Omni prompt
fal's Omni 1.1 example ends its prompt with a quoted spoken line. Here is that pattern on Sume from a still, the cost of 6 seconds, and a transcript check.
- Photo to watercolor with ChatGPT Image 2: one reference, one prompt
The Sume docs use ChatGPT Image 2 with one reference photo and the prompt "make this scene look like a watercolor painting". The request and what to check.
- Photographer highlight reel: trim three clips, render one timeline
A freelance photographer or videographer can cut three client clips with video-trim and join them with a music bed in Timeline 1.0. Rates, limits, and rights.
- Pick clip lengths from a voice-over script: one sentence per clip
Turn a voice-over into clips: measure each voiced sentence, round up, and pick the models whose duration window contains it. Windows for six Sume models.
Written by Sume