Dentist office welcome video: a 9:16 AI avatar clip from one script
A dental practice can make a vertical welcome video without a shoot: one avatar, one script of 4 to 60 seconds, burned-in captions, and a staff review step.
A dental office can make a welcome video without a camera crew by creating one avatar, writing one script, and calling the avatar talking-video route. Sume accepts a script when its estimated length is 4 to 60 seconds, outputs 9:16 by default, and can burn captions into the finished MP4.
This post covers the practical path for a front-desk manager or a small practice owner. It also covers where Sume stops: it does not check clinical or advertising rules for your region, so a person at the practice must read the script and the result before it goes on the website or a reel.
Why a welcome clip fits the trend
A roundup of this month's AI video trends lists avatars as common in training, support, and localized sales material, and calls 5 to 30 seconds the practical range for polished short-form clips (AI Video Generation Trends, read 2026-10-07). A welcome message from the front desk is exactly that kind of material: one presenter, one message, reused for months.
The rest of the work is small. You make the avatar once, then every new message (new hours, a new hygienist, a holiday closure) is a new script against the same avatar handle.
Make the presenter once
Avatar creation takes a top-level avatar_handle and an input that is a prompt, a profile of traits, or a reference image. The job returns a handle you reuse. Wait for the job to complete before you generate videos with it, because the talking-video route needs a ready avatar.
Pick a neutral, friendly presenter and keep the handle in your notes. Do not build the avatar to look like a named real person unless you have that person's consent; Sume's docs do not make that judgment for you.
curl -X POST https://api.sume.com/v1/avatar-1.0/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: dental-host-001" \
-d '{
"avatar_handle": "front_desk_host",
"input": { "type": "props", "sex": "female", "age": 34, "ethnicity": "Asian" }
}'Write a script that fits the window
Sume plans the clip from your script, and a script it estimates at under 4 or over 60 seconds is rejected. A welcome message of 70 to 110 spoken words usually lands inside the window; if you have more to say, split it into two jobs rather than rushing the read.
Send aspect_ratio as 9:16 for a phone-first clip, or 16:9 for the website hero. Resolution is 720p at this time, and quality accepts standard, plus (the default), and max.
| Setting | Value | What it means for the office |
|---|---|---|
| Script length | 4 to 60 seconds, estimated | Split longer messages into separate jobs |
| aspect_ratio | 1:1, 3:4, 9:16, 4:3, 16:9; default 9:16 | 9:16 for reels, 16:9 for a site banner |
| resolution | 720p | Fine for phones; plan for it on a big screen |
| quality | standard, plus (default), max | standard is the fastest path, max is slower |
| captions | Optional inline, style slam by default | Useful because many viewers watch muted |
Generate, caption, and review
Send the script with avatar_handle and turn on inline captions. Inline captions are limited to videos Sume estimates at 60 seconds or less, and a caption failure is a soft failure: the avatar job can still succeed with a clean video and captions.status=failed. Poll the job and fetch the result as described in the jobs guide.
Before publishing, have someone at the practice check the spoken text against your real hours, services and any wording your licensing body or ad rules care about. The API produces a video; it does not certify claims.
curl -X POST https://api.sume.com/v1/avatar-1.0/talking-video \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: dental-welcome-001" \
-d '{
"avatar_handle": "front_desk_host",
"aspect_ratio": "9:16",
"quality": "plus",
"script": "Welcome to our office. We are open Monday to Friday, and new patients can book a first visit online. Our team will walk you through every step.",
"captions": { "enabled": true, "style": "slam", "language": "auto" }
}'What this does not cover
The call returns a file on media.sume.com. Posting it to your site, Google profile or social account is a separate step you do yourself. If you want to preview the first frames before paying for a full render, the avatar video preview route exists for that, and it is documented separately.
Sources
Related posts
More in Use cases
- Destination film from ten location photos with Wan 3.0 on Sume
Turn ten photos of a town into a 30-second film with wan-3.0 reference images: $3.75 at 720p, $7.50 at 1080p, and a prompt that needs a dawn-to-dusk arc.
- Dictation audio: 12 sentences in one TTS job, 6 cents, wav slices
Make 12 dictation sentences as separate wav files from one Sume TTS job with sentence segmentation. 1,080 characters cost 6 cents, against 12 cents as 12 jobs.
- Dolly zoom (vertigo effect) prompt for Gemini Omni, tested cheap
Ask Gemini Omni for a dolly zoom by describing both motions, try it at 360p for 30 cents, and check the background stretch before paying for 1080p on Sume.
- Edit once and crop three ways, or edit three times: 4:5, 1:1, 9:16
Crop one 4:5 edit to 1:1 and 9:16 in Pillow, or edit per shape. A crop keeps 80% of the height or 70.3% of the width; three edits cost three times as much.
Written by Sume