Tutor intro video: approve the first frame, then pick the tier
A tutor can approve an avatar intro's first frame before paying for the full render, then choose standard, plus or max at generate-video without a new preview.
To make a tutor intro video you can check before you pay for it, create an avatar video preview, review the first-frame still, and call generate-video on the preview id. The preview stills are independent of the quality tier, so you can approve once and choose standard, plus or max for the final render without making a new preview.
The flow is from Avatar video previews and Generate avatar video, read 2026-10-04.
One avatar, many intros
Avatar 1.0 is two steps: create a reusable avatar, then use its handle to make talking videos. A tutor can build the avatar from a photo with input.type: "photo", a prompt, or structured traits, and keep the handle for every later intro.
Create the preview
The create body matches the Avatar Video fields: exactly one of script or video_inputs, plus optional scene, quality, aspect_ratio, title and captions. Captions you set here are stored and burned in only when you generate the video, because preview stills are never caption-burned.
curl -X POST https://api.sume.com/v1/avatar-video-previews \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: tutor-intro-preview-001" \
-d '{
"avatar_handle": "math_tutor",
"script": "Hi, I am Ms. Rivera. In our first session we will find where algebra stops making sense, and fix that first.",
"scene": { "type": "prompt", "prompt": "A bright home study with a whiteboard behind the tutor" },
"aspect_ratio": "16:9",
"quality": "standard",
"captions": { "enabled": true, "style": "slam", "language": "auto" }
}'Review, regenerate, generate
The response includes job polling URLs and an avatar_video_preview_id. When the preview is ready, GET /v1/avatar-video-previews/:id returns preview_image_url, and scene_previews[] for multi-scene plans. If the framing is wrong, POST /v1/avatar-video-previews/:id/regenerate reuses the stored request and refreshes only the stills.
| Call | Changes | Keeps |
|---|---|---|
| POST /v1/avatar-video-previews | Makes first-frame stills | Nothing yet |
| POST .../:id/regenerate | New stills | Avatar, script, scene, tier, aspect ratio |
| POST .../:id/generate-video | Starts the final video | The approved first frame |
| generate-video with quality | Final render tier only | Preview stills |
The limits that still apply
Final render: curl -X POST https://api.sume.com/v1/avatar-video-previews/$PREVIEW_ID/generate-video with a body of {"quality": "max"}, or an empty body to keep the tier chosen at create. Admission, the ledger reservation and the provider submit all use that effective tier.
Structural fields (script, video_inputs, avatar_handle, scene, aspect_ratio) cannot be overridden at this step. Changing the script means a new preview. The video must still estimate to 4 to 60 seconds.
Sources
Related posts
More in Use cases
- 2-second YouTube intro sting: Wan 3.0 has the lowest floor
The shortest clip length Sume's docs list for the Video Router is 2 seconds on wan-3.0. Seedance 2.5 starts at 4. Here is a request and how to trim for less.
- UGC-style ad on Sume: a Format run or the avatar endpoint
Two ways to make a UGC-style ad on Sume: a catalog Format run for a full cut, or the avatar talking-video endpoint for a 4 to 60 second presenter clip.
- Vacation rental check-in video: an avatar host with silence beats
One avatar, one shared background and silence beats let a host walk guests through check-in in 60 seconds or less. Set video_inputs on Avatar Video.
- Vendor pages to read before AI music goes in an ad
The vendor pages that answer commercial-use questions for Suno v6, ElevenLabs Music v2.5 and Google Lyria 3.5, and which question each answers, read 2026-10-04.
Written by Sume