AI training video generator: from written steps to video

An AI training video generator turns written steps into presenter videos. Build one module from short avatar clips: one clip per step, captioned, joined.

5 min readSume
All posts

An AI training video generator turns a written procedure or script into a video in which an AI presenter explains each step, so a team can make and update training without filming. One way to build it: split the material into steps, render one presenter clip per step, caption each clip, then join the clips into one module or publish them one by one.

On Sume that means one reusable avatar, one talking video per step, a caption job per clip, and a Timeline render to join them. Facts come from Generate avatar video, Video captions and Timeline 1.0, read on 2026-09-28, plus Sume's current code where noted. Selling a course rather than training staff? AI avatar for online course videos covers that case.

How do I turn a procedure into a training video?

  • Script each step the way a trainer would say it: one action per step, in the order the trainee does them.
  • Write one step per clip. A talking video is accepted when Sume estimates it at 4-60 seconds, so a step that runs longer becomes two clips.
  • Create one avatar and keep its handle. The docs recommend a stable handle so every step reuses the same presenter; creating a reusable AI avatar shows the call.
  • Render each step with the same avatar_handle, aspect_ratio and quality. 16:9 is the landscape option; the default is 9:16.
  • Poll each job; completed results can include a public media.sume.com video.
curl -X POST https://api.sume.com/v1/avatar-1.0/talking-video \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: safety-module-step-03" \
  -d '{
    "avatar_handle": "training_host",
    "script": "Step three: lock the machine before you open the guard.",
    "aspect_ratio": "16:9",
    "quality": "plus"
  }'

How do I caption the steps and join them into one module?

Caption each step clip before any join. In current code the caption job (POST /v1/video-captions) refuses a source longer than 60 seconds or one with no audio stream, so a finished module can't be captioned in one go, while a single step fits. Burning captions onto a video covers the request.

To join, one Timeline 1.0 render takes the step clips as this workspace's media.sume.com files and returns one MP4 of up to 1,800 seconds. In current code the render drops each clip's own audio, so the steps' speech has to go on its audio spine; making an avatar video longer than 60 seconds walks through that join.

Should a training module be one video or separate steps?

Procedures change, so plan for updates before you pick:

  • Separate step videos: when one step changes, you render one new clip and swap it in. The other steps stay as they are and are not billed again.
  • One joined module: easier to assign as a single video, but a changed step means a new clip plus a new Timeline render of the whole module, billed by output minute.
  • Either way, send a changed script with a new Idempotency-Key. Sume answers 409 idempotency_conflict when a key is reused for a different payload, so reuse a key only for an exact retry.

What can't it do?

  • Speak other languages from the avatar route: current code makes Avatar 1.0 talking videos speak English only. What languages can an AI avatar speak covers the TTS and lip sync route for other languages.
  • Use screen recordings: there is no public upload route for local files, so software walkthroughs need another tool.
  • Put two presenters in one clip: current execution supports one resolved avatar per video.
  • Go above 720p: resolution is currently 720p.

What does an AI training video cost?

Each piece bills separately, plus a 5.5% agent fee by default:

From Generate avatar video, Video captions, Timeline 1.0 and API pricing, read 2026-09-28.
PieceLimitPrice
AvatarCreated once, reused by handle$0.95 per avatar
Step clipEstimated 4-60 seconds each$0.184/s standard, $0.245/s plus, $0.55/s max (no product image)
Caption jobOne clip up to 60 secondsFixed per clip; live price in GET /v1/catalog
Timeline joinUp to 1,800 seconds$0.10 per output minute

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume