Customer service training videos: role-plays made with AI
Customer service training videos as short AI role-plays: which scenarios to script, how to build a two-person scene, captions, and what each clip costs.

A useful customer service training video is a short role-play: a customer says something difficult, the employee answers the way you want it handled, and a narrator names the skill being shown. You can make a library of these with AI presenters instead of actors, one scenario per video, each under a minute so staff can watch one before a shift.
The Sume facts below come from the Generate avatar video, Create new avatar, Audio detach, Timeline 1.0 and Video captions docs, read on 2026-09-29. Points marked as current behavior are read from Sume's code.
Which customer service scenarios should I script?
Pick the conversations your team gets wrong or finds hard, and write each as three to six short turns plus one line from the narrator. Examples:
- An upset customer whose order is late: acknowledge, give a fact, give a next step.
- A refund request outside your policy: say no without arguing, offer what you can do.
- A question the employee can't answer: the hand-off to someone who can.
- Retail: a return at the counter without a receipt.
- Healthcare front desk: rescheduling a missed appointment. Use invented names and details, never real patient information.
- The same scene twice, handled badly and then well, if your trainers teach by contrast.
How do I make a two-person role-play with AI?
Sume's current execution supports one resolved avatar per final video, so the customer and the employee never share a shot. Create one presenter per part once (customer, employee, and a narrator if you use one), each from a prompt, profile traits or a reference photo, and reuse the same cast across every scenario so staff recognize the roles. Render one talking clip per turn. Each clip needs a script Sume estimates at 4 to 60 seconds, so give every reply enough words to fill four seconds; a one-word answer on its own falls short.
Then cut the turns together in speaking order in one Timeline render. In current code the render drops each clip's own sound, so each turn's voice is detached first and laid on the render's audio spine. AI avatar conversation video: two speakers shows that join step by step.
When a policy changes, re-render only the turns whose words changed and run the join again; the presenters and the other turns stay as they are.
Should training videos have captions?
Caption them, so staff can watch on a shop floor or with the sound off. Send the finished scene to POST /v1/video-captions; script_text aligns the burned words to your script. In current code the caption job refuses a video longer than 60 seconds or one without an audio stream, which is one more reason to keep each scenario under a minute.
How much do customer service training videos cost?
Each step is billed per call from one prepaid balance. The presenters are a one-time cost; the clips are billed per second, so a scene's price follows its length.
| Step | Call | Price |
|---|---|---|
| Presenters: customer, employee, narrator (once) | POST /v1/avatar-1.0/generate | $0.95 per avatar, each |
| One clip per turn, default quality | POST /v1/avatar-1.0/talking-video | $0.245 per second |
| Pull each turn's voice | POST /v1/audio-detach | $0.01 per job |
| Cut the turns together | POST /v1/timeline-1.0/render | $0.10 per output minute |
| Burned-in captions | POST /v1/video-captions | $0.20 per job, for videos up to 60 seconds |
What are the limits?
- No shared shot: each clip holds one presenter.
- English only: in current code the avatar route speaks English. What languages can an AI avatar speak covers the route for other languages.
- Up to 20 turns per joined scene when the voices go in as
audio.parts[]. - The render reads only this workspace's
media.sume.comfiles, such as the clips and audio Sume returned; don't plan on uploading recordings of real calls.
Sources
Related posts
More in Use cases
- Dental marketing videos with AI: what to make, what to skip
Dental marketing videos AI can make without a patient: a meet-the-dentist clip, first-visit and FAQ answers, and office-photo tours. What to skip, and costs.
- Donor thank you video: one render, each donor's name
A donor thank you video can show each donor's name for the price of one render plus a caption job per name, or say each name aloud at a full render per donor.
- Financial advisor video marketing with AI, compliance first
Financial advisor video marketing with AI: short explainers from an avatar, scripts approved before rendering, exact on-screen text, and per-second costs.
- Gym promo video with AI: from your own gym photos
Make a gym promo video without a shoot: animate photos of your own floor and classes, add a voiced offer, music and captions, and render a vertical cut.
Written by Sume