FAQ videos: how to make them with an AI avatar
A FAQ video answers one common question in a short clip. How to make a set with an AI avatar: one script, one render, and one presenter throughout.
A FAQ video answers one frequently asked question in a short clip, so a customer can watch the answer instead of reading it, in a help center, an onboarding email or a support reply. With an AI avatar you write each answer as a script and render one talking video per question: the same presenter every time, and no filming.
On Sume, each answer is an Avatar 1.0 talking video of up to 60 seconds, about 165 words by Sume's current estimate. The videos are pre-rendered answers; the avatar doesn't answer customers live. The Sume facts come from the Generate avatar video and Avatar video previews docs and the Sume API reference, read on 2026-09-28. Anything called current behavior is read from Sume's code.
What makes a good FAQ video?
- One question per video, named in the first line, so the clip still makes sense when someone lands on it from search or a chat link.
- The answer first, then the steps. Leave out anything that belongs to another question.
- Under a minute. Longer topics become two questions.
- Captions, so the answer reads with the sound off.
- The same presenter, framing and shape across the set, so it looks like one library.
How do I make FAQ videos with an AI avatar?
Create one avatar and reuse its handle for every answer; the docs suggest a stable handle, a simple name your app reuses. Then send one POST /v1/avatar-1.0/talking-video per question, with a script Sume estimates at 4–60 seconds; Script length for an AI avatar video explains how the words are counted.
Each video is a job: submit it, poll the status, then read the result, which can include public media.sume.com video artifacts. Give every answer its own Idempotency-Key, and reuse a key only for the same payload.
curl -X POST https://api.sume.com/v1/avatar-1.0/talking-video \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: faq-reset-password-v1" \
-d '{
"avatar_handle": "acme_support",
"script": "Forgot your password? On the sign-in page, choose Forgot password, and we will email you a link to set a new one.",
"aspect_ratio": "16:9",
"title": "How do I reset my password?"
}'How do I add captions to FAQ videos?
Not on the request above: in current code POST /v1/avatar-1.0/talking-video refuses captions with 400 invalid_request, even though the reference lists it. Use one of two routes instead:
- Create an avatar video preview with
captionson it. The preview stores them and burns them in only when you callgenerate-video, and you can check the first frame on the way. - Caption the finished answer with
POST /v1/video-captions. In current code that job refuses a source longer than 60 seconds or one with no audio stream.
How do I update one answer?
Render that answer again and leave the rest alone. The words, voice and scene are part of each render, so a changed answer is a new job with the same avatar handle, a new script and a new Idempotency-Key. Then point the help-center article or saved reply at the new URL.
Can an AI avatar answer customer questions live?
Not a Sume avatar. Each video is rendered as a job, and even mode: sync waits at most 30 seconds; that ceiling bounds the HTTP wait, not the job. What a support team can do is link the rendered answer wherever the question comes up, and keep the scripts accurate as the product changes.
What are the limits, and what does it cost?
Creating the avatar costs $0.95 per avatar, once, and each answer is billed per second by quality tier, $0.184/s standard, $0.245/s plus, $0.55/s max (no product image); plus is the default tier. Both are plus a 5.5% agent fee by default, and API pricing lists every rate. Captions add to the bill: stored on a preview they are a caption add-on to that video, and a standalone caption job reserves a fixed amount per clip of up to 60 seconds, with live prices in GET /v1/catalog.
The talking video speaks English only in current code; Which languages can an AI avatar speak? covers answers in other languages.
| Each FAQ answer | |
|---|---|
| Length | Estimated 4–60 seconds |
| Presenter | One avatar per video; reuse one handle across the set |
| Shape | 1:1, 3:4, 9:16, 4:3, 16:9; default 9:16 |
| Resolution | 720p is the documented resolution |
| Speech | English only |
| Captions | Through a preview, or a caption job on a clip of 60 seconds or less with audio |
| Result | Public media.sume.com video artifacts |
Sources
Related posts
More in Use cases
- Can you add videos to Google Business Profile? Rules and tips
Yes. Google Business Profile takes videos up to 30 seconds, up to 75 MB, at 720p or higher, reviewed before they show. How to add one and make one.
- How to make a Slack emoji with AI (under 128 KB)
Slack says square images under 128 KB with transparent backgrounds work best. Generate one, cut it out to a PNG, shrink it, then click Add Emoji.
- How to make a YouTube channel trailer with AI
A YouTube channel trailer greets people who haven't subscribed. What to say, how long to make it, and how to render one with an AI avatar.
- How to make AI ASMR videos: prompt the picture and the sound
Use a video model that generates sound with the picture, and describe the close-up action and its sounds in the prompt. Models, prompts, and length.
Written by Sume