Personal trainer class promo: avatar hook, silent demo beat, CTA

Build an 11-second class promo for a fitness studio from three ordered scenes on one avatar, with a silence beat for the demo. Limits and request body included.

4 min readSume
All posts

A personal trainer or small studio can make a class promo as three ordered scenes on one avatar: a spoken hook, a silent beat for the demo, and a spoken call to action. Sume's avatar talking-video route takes ordered video_inputs for this, with a planned total of 4 to 60 seconds.

The silent beat is the useful trick. A voice.type of silence with a duration gives you a gap with no speech where you can lay your own gym footage later, without the avatar talking over it.

Why scenes beat one long script

This month's trend roundups say avatars are mainly used for training, support and localized sales material, and that agents are being used to draft scripts and variants (AI Video Generation Trends, read 2026-10-07). A class promo is a small version of the same job: a repeatable structure with new words each week.

A single script gives you one continuous read. video_inputs lets you set the length of each beat, so the hook stays under four seconds and the call to action gets the room it needs.

The request

Create the avatar once with POST /v1/avatar-1.0/generate, then reuse its handle. In the talking-video body, send either script or video_inputs, never both. Spoken scenes use voice.type: "text" with one of script or input_text. A silent scene needs a duration and no text.

The current execution supports one resolved avatar per final video, and expects the scene backgrounds to resolve to one shared scene, so keep the background prompt the same across scenes.

curl -X POST https://api.sume.com/v1/avatar-1.0/talking-video \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: spin-class-promo-001" \
  -d '{
    "avatar_handle": "studio_coach",
    "aspect_ratio": "9:16",
    "quality": "plus",
    "video_inputs": [
      { "id": "hook", "voice": { "type": "text", "script": "Thursday 6 a.m. spin is back.", "duration": 3 },
        "background": { "type": "prompt", "prompt": "Bright studio, mirrors, soft morning light" } },
      { "id": "demo", "voice": { "type": "silence", "duration": 4 },
        "background": { "type": "prompt", "prompt": "Bright studio, mirrors, soft morning light" } },
      { "id": "cta", "voice": { "type": "text", "script": "Book your first class free from the link below.", "duration": 4 },
        "background": { "type": "prompt", "prompt": "Bright studio, mirrors, soft morning light" } }
    ]
  }'

Pick the quality tier

Quality is standard, plus or max. plus is the default and the balanced path, standard is the fastest Sume path, and max is the highest tier with slower turnaround. For a weekly promo that a trainer will post and replace, standard or plus is usually the practical choice.

Poll the job at /v1/jobs/{id}/status and fetch /v1/jobs/{id}/result when it is ready.

Scene-based avatar video rules (Sume docs, read 2026-10-07)
RuleValueWhy it matters for a class promo
Total planned length4 to 60 secondsHook + demo + CTA fits easily
script vs video_inputsSend one, not bothUse video_inputs for beats
Silence scenevoice.type silence plus durationRoom for your own footage
Avatars per videoOne resolved avatarKeep one coach
qualitystandard, plus (default), maxplus is the balanced default

Where to add the real footage

Sume does not film your class. To put real class footage under the silent beat, export the avatar clip and edit it in your own editor, or assemble clips with Sume's timeline render. Disclose that the presenter is an AI avatar where your market or the platform expects it.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume