AI training video generator: from written steps to video
An AI training video generator turns written steps into presenter videos. Build one module from short avatar clips: one clip per step, captioned, joined.

An AI training video generator turns a written procedure or script into a video in which an AI presenter explains each step, so a team can make and update training without filming. One way to build it: split the material into steps, render one presenter clip per step, caption each clip, then join the clips into one module or publish them one by one.
On Sume that means one reusable avatar, one talking video per step, a caption job per clip, and a Timeline render to join them. Facts come from Generate avatar video, Video captions and Timeline 1.0, read on 2026-09-28, plus Sume's current code where noted. Selling a course rather than training staff? AI avatar for online course videos covers that case.
How do I turn a procedure into a training video?
- Script each step the way a trainer would say it: one action per step, in the order the trainee does them.
- Write one step per clip. A talking video is accepted when Sume estimates it at 4-60 seconds, so a step that runs longer becomes two clips.
- Create one avatar and keep its handle. The docs recommend a stable handle so every step reuses the same presenter; creating a reusable AI avatar shows the call.
- Render each step with the same
avatar_handle,aspect_ratioandquality.16:9is the landscape option; the default is9:16. - Poll each job; completed results can include a public
media.sume.comvideo.
curl -X POST https://api.sume.com/v1/avatar-1.0/talking-video \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: safety-module-step-03" \
-d '{
"avatar_handle": "training_host",
"script": "Step three: lock the machine before you open the guard.",
"aspect_ratio": "16:9",
"quality": "plus"
}'How do I caption the steps and join them into one module?
Caption each step clip before any join. In current code the caption job (POST /v1/video-captions) refuses a source longer than 60 seconds or one with no audio stream, so a finished module can't be captioned in one go, while a single step fits. Burning captions onto a video covers the request.
To join, one Timeline 1.0 render takes the step clips as this workspace's media.sume.com files and returns one MP4 of up to 1,800 seconds. In current code the render drops each clip's own audio, so the steps' speech has to go on its audio spine; making an avatar video longer than 60 seconds walks through that join.
Should a training module be one video or separate steps?
Procedures change, so plan for updates before you pick:
- Separate step videos: when one step changes, you render one new clip and swap it in. The other steps stay as they are and are not billed again.
- One joined module: easier to assign as a single video, but a changed step means a new clip plus a new Timeline render of the whole module, billed by output minute.
- Either way, send a changed script with a new
Idempotency-Key. Sume answers409 idempotency_conflictwhen a key is reused for a different payload, so reuse a key only for an exact retry.
What can't it do?
- Speak other languages from the avatar route: current code makes Avatar 1.0 talking videos speak English only. What languages can an AI avatar speak covers the TTS and lip sync route for other languages.
- Use screen recordings: there is no public upload route for local files, so software walkthroughs need another tool.
- Put two presenters in one clip: current execution supports one resolved avatar per video.
- Go above 720p:
resolutionis currently720p.
What does an AI training video cost?
Each piece bills separately, plus a 5.5% agent fee by default:
| Piece | Limit | Price |
|---|---|---|
| Avatar | Created once, reused by handle | $0.95 per avatar |
| Step clip | Estimated 4-60 seconds each | $0.184/s standard, $0.245/s plus, $0.55/s max (no product image) |
| Caption job | One clip up to 60 seconds | Fixed per clip; live price in GET /v1/catalog |
| Timeline join | Up to 1,800 seconds | $0.10 per output minute |
Sources
Related posts
More in Use cases
- Arcads alternatives: AI UGC ad tools by plan and API
Arcads alternatives for AI-actor UGC ads: Creatify, HeyGen, Higgsfield and Sume compared on presenters, API, billing and free options, read 2026-09-28.
- Arcads vs Creatify: AI actor ads or URL-to-video ads
Arcads builds ads from your script and 1,000+ AI actors on monthly plan credits. Creatify turns a product URL into a video ad; API plans start at $99.
- Can you schedule YouTube Shorts? Yes, in Studio or by API
Yes. Upload a Short in YouTube Studio and set a publish time on the Schedule card, or set status.publishAt on a private upload with the Data API.
- CGI product videos with AI: the 3D ad look from a photo
AI can imitate the CGI product-ad look from a packshot and a scene prompt. You get a generated video clip, not a 3D model you can re-render.
Written by Sume