Driver onboarding videos with an AI avatar: five 30-second clips
Five short avatar clips can cover a driver's first day. What each should say, how to handle a multilingual crew honestly, and the cost by tier.
Delivery and gig drivers need a few instructions on day one: how to start a shift, how to hand over a parcel, what to do when something goes wrong, how pay is shown, and where to get help. Five 30-second avatar clips cover that in under three minutes of viewing, rendered once and reused for every new hire. They do not replace a safety briefing or a ride-along.
Five clips, one topic each
- Start of shift: opening the app, checking the vehicle, where the shift ends.
- Handover: confirmation steps and what counts as proof of delivery.
- Problems: damaged parcel, missing address, unsafe location. Name the person or number to contact in the page text.
- Pay: when and where earnings are displayed. No figures in the video; they change.
- Help: how to reach support, and what to prepare before you call.
The language question
Sume's Avatar 1.0 speaks English only. If your crew reads mostly another language, a spoken-English clip is not enough on its own. You can burn subtitle text in another language onto the finished clip with the standalone captions endpoint using authored cues (text, start, end), which skips speech-to-text, but you must write and time those cues yourself, and a native speaker should check them. For safety-critical content, a translated clip recorded by a person is a better choice than subtitles.
Cost
Using one avatar (0.95 dollars once) and five clips of 30 seconds each with no product image:
| Tier | Per clip | Five clips | With avatar creation |
|---|---|---|---|
| standard | $5.52 | $27.60 | $28.55 |
| plus (default) | $7.35 | $36.75 | $37.70 |
| max | $16.50 | $82.50 | $83.45 |
Keep them current
Because the clip is an MP4, an app screen change makes it wrong. Keep the script in version control, note the app version it was written for, and re-render the one clip that changed. Rendering one 30-second clip again costs the same as the first time, and the other four stay valid.
Use a plain, neutral scene prompt, such as a depot office, so the background does not suggest a specific vehicle or brand you do not own. A scene photo of your own depot is another option and must be a public HTTPS image URL.
Review before launch
Have an operations lead read the five scripts. Check each finished transcript against the approved text, since a mis-rendered number or name in a safety or pay clip causes real confusion. Then publish them in your onboarding page in order with a short text summary under each.
Measuring whether it works
Track simple things: whether new drivers complete the five clips, and which topic produces the most support contacts in week one. If handover errors stay high, rewrite that script rather than adding length. Because each clip is its own job, a rewrite and re-render of one 30-second clip is 7.35 dollars at plus. You can also compare two script variants for the same topic by rendering both.
Sources
Related posts
More in Use cases
- Destination film from ten location photos with Wan 3.0 on Sume
Turn ten photos of a town into a 30-second film with wan-3.0 reference images: $3.75 at 720p, $7.50 at 1080p, and a prompt that needs a dawn-to-dusk arc.
- Dictation audio: 12 sentences in one TTS job, 6 cents, wav slices
Make 12 dictation sentences as separate wav files from one Sume TTS job with sentence segmentation. 1,080 characters cost 6 cents, against 12 cents as 12 jobs.
- Edit once and crop three ways, or edit three times: 4:5, 1:1, 9:16
Crop one 4:5 edit to 1:1 and 9:16 in Pillow, or edit per shape. A crop keeps 80% of the height or 70.3% of the width; three edits cost three times as much.
- How do I make AI announcer intros for conference speakers?
Ten 25-second speaker intros with one shared sting cost about 43 cents on Sume: 2 cents of TTS and a 1-cent join each, plus a $0.125 track made once.
Written by Sume