AI Avatar Welcome Video for New Dental Patients: Script and Cost
A scripted avatar welcome clip for a dental practice: scene plan, word-to-second math, request JSON and cost on each Sume Avatar 1.0 quality tier.
A dental practice can send new patients a short welcome video from a reusable avatar host: who they will meet, where to park, what to bring. With Sume Avatar 1.0 that is a scripted render of 4 to 60 seconds, not a live chat, so the script is where the work is.
Prices are from Sume's API pricing page and the request shape from Generate avatar video, both read on 2026-10-05. The avatar must exist first: Create new avatar accepts a prompt, a profile or a photo.
What should the script and scenes look like?
Three scenes keep a first-visit clip tidy: greeting, what to bring, and a reassurance beat. Sume estimates a spoken scene at ceil(words / 2.8) seconds, so the counts below are the planning numbers, not a measured render time.
The seconds column is Sume's planning estimate of ceil(words / 2.8) per scene, not a measured render time.
| Scene | Spoken line | Words | Seconds |
|---|---|---|---|
| greet | "Welcome to Bright Row Dental. I'm your virtual h..." | 18 | 7 |
| bring | "Please bring your photo ID, your insurance card,..." | 16 | 6 |
| calm | "Arrive ten minutes early, and we'll take it slow..." | 14 | 5 |
| Total | 48 | 18 |
POST https://api.sume.com/v1/avatar-1.0/talking-video
{
"avatar_handle": "front_desk_host",
"aspect_ratio": "9:16",
"quality": "plus",
"video_inputs": [
{ "id": "greet", "voice": { "type": "text", "script": "Welcome to Bright Row Dental. I'm your virtual host, and this is what your first visit looks like." }, "background": { "type": "prompt", "prompt": "Bright, plain room, locked-off camera, soft daylight" } },
{ "id": "bring", "voice": { "type": "text", "script": "Please bring your photo ID, your insurance card, and a list of any medications you take." }, "background": { "type": "prompt", "prompt": "Bright, plain room, locked-off camera, soft daylight" } },
{ "id": "calm", "voice": { "type": "text", "script": "Arrive ten minutes early, and we'll take it slowly from there. See you soon." }, "background": { "type": "prompt", "prompt": "Bright, plain room, locked-off camera, soft daylight" } }
]
}What does it cost?
The plan above comes to 18 seconds, inside the 4 to 60 second window. At Sume's per-second rates (no product image), one render costs $3.31 on standard, $4.41 on plus (the default when quality is omitted) and $9.90 on max. Across 40 patient intake emails, that is $132.48, $176.40 and $396.00. Creating the avatar is a separate one-time $0.95; later clips reuse the handle.
| Quality | Rate per second | One clip | 40 patient intake emails |
|---|---|---|---|
| standard | $0.184 | $3.31 | $132.48 |
| plus | $0.245 | $4.41 | $176.40 |
| max | $0.55 | $9.90 | $396.00 |
How do you keep a health-adjacent script safe?
Keep the avatar to logistics. Parking, forms and arrival time are things a script can state exactly; clinical advice is not. Write each fact into the script yourself and review the rendered clip before it goes out.
Render one clip first. Avatar video previews produce a still of the first frame, so you can check framing and the avatar's look before paying for the full video.
What will it not do?
It will not answer a patient's questions. The clip is the same for everyone who receives it, and Sume does not claim real-time conversation. If a patient needs to ask something, link them to your phone line or booking page.
Scripts for Avatar Video should be English; Sume's guidance is that non-English speech breaks caption alignment.
Sources
Related posts
More in Use cases
- AI backdrops for TikTok Shop: staging must match what you sell
TikTok Shop's page says LIVE backgrounds and staging items must match the products sold. How to review a generated backdrop and use video filter dim or crop.
- AI music as the main focus of a Short: YouTube says to disclose it
YouTube lists music that is the main focus of a video as altered or synthetic content to disclose. How that applies to a Short built on a Sume music track.
- All-hands recap with an AI presenter: a 60-second video and its cost
A weekly company recap read by a Sume avatar: script structure, the 60-second cap, a cost table for one week and a year of 52, and the disclosure to add.
- Amazon will flag AI-generated people to shoppers: what to plan
Amazon says it will add a shopper indicator where an image includes AI-generated people, in all stores. Here is what is known, what is open, and how to plan.
Written by Sume