AI Avatar Welcome Video for New Dental Patients: Script and Cost

A scripted avatar welcome clip for a dental practice: scene plan, word-to-second math, request JSON and cost on each Sume Avatar 1.0 quality tier.

4 min readSume
All posts

A dental practice can send new patients a short welcome video from a reusable avatar host: who they will meet, where to park, what to bring. With Sume Avatar 1.0 that is a scripted render of 4 to 60 seconds, not a live chat, so the script is where the work is.

Prices are from Sume's API pricing page and the request shape from Generate avatar video, both read on 2026-10-05. The avatar must exist first: Create new avatar accepts a prompt, a profile or a photo.

What should the script and scenes look like?

Three scenes keep a first-visit clip tidy: greeting, what to bring, and a reassurance beat. Sume estimates a spoken scene at ceil(words / 2.8) seconds, so the counts below are the planning numbers, not a measured render time.

The seconds column is Sume's planning estimate of ceil(words / 2.8) per scene, not a measured render time.

Scene plan and estimated seconds (read 2026-10-05)
SceneSpoken lineWordsSeconds
greet"Welcome to Bright Row Dental. I'm your virtual h..."187
bring"Please bring your photo ID, your insurance card,..."166
calm"Arrive ten minutes early, and we'll take it slow..."145
Total4818
POST https://api.sume.com/v1/avatar-1.0/talking-video
{
  "avatar_handle": "front_desk_host",
  "aspect_ratio": "9:16",
  "quality": "plus",
  "video_inputs": [
    { "id": "greet", "voice": { "type": "text", "script": "Welcome to Bright Row Dental. I'm your virtual host, and this is what your first visit looks like." }, "background": { "type": "prompt", "prompt": "Bright, plain room, locked-off camera, soft daylight" } },
    { "id": "bring", "voice": { "type": "text", "script": "Please bring your photo ID, your insurance card, and a list of any medications you take." }, "background": { "type": "prompt", "prompt": "Bright, plain room, locked-off camera, soft daylight" } },
    { "id": "calm", "voice": { "type": "text", "script": "Arrive ten minutes early, and we'll take it slowly from there. See you soon." }, "background": { "type": "prompt", "prompt": "Bright, plain room, locked-off camera, soft daylight" } }
  ]
}

What does it cost?

The plan above comes to 18 seconds, inside the 4 to 60 second window. At Sume's per-second rates (no product image), one render costs $3.31 on standard, $4.41 on plus (the default when quality is omitted) and $9.90 on max. Across 40 patient intake emails, that is $132.48, $176.40 and $396.00. Creating the avatar is a separate one-time $0.95; later clips reuse the handle.

Avatar Video, no product image, 18 seconds, 40 patient intake emails (rates read 2026-10-05)
QualityRate per secondOne clip40 patient intake emails
standard$0.184$3.31$132.48
plus$0.245$4.41$176.40
max$0.55$9.90$396.00

How do you keep a health-adjacent script safe?

Keep the avatar to logistics. Parking, forms and arrival time are things a script can state exactly; clinical advice is not. Write each fact into the script yourself and review the rendered clip before it goes out.

Render one clip first. Avatar video previews produce a still of the first frame, so you can check framing and the avatar's look before paying for the full video.

What will it not do?

It will not answer a patient's questions. The clip is the same for everyone who receives it, and Sume does not claim real-time conversation. If a patient needs to ask something, link them to your phone line or booking page.

Scripts for Avatar Video should be English; Sume's guidance is that non-English speech breaks caption alignment.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume