AI avatar for healthcare: patient education videos

Use an AI avatar in healthcare for general patient education: short, clinician-reviewed videos, captioned for waiting rooms, with no patient data.

5 min readSume
All posts

An AI avatar in healthcare fits recorded, general patient education: how to prepare for a visit, aftercare steps, clinic information, or a waiting-room explainer. The avatar speaks a script your clinicians wrote or reviewed, so keep each video to one topic, keep patient-specific information out of it, and caption it for rooms where the sound is off. On Sume, each topic is an Avatar 1.0 talking video of up to 60 seconds, and other languages go through TTS 1.0 plus lip sync.

Sume facts come from the Generate avatar video and Avatar video previews docs and the Sume API reference, read on 2026-09-28; anything called current behavior is read from Sume's code. This is a production workflow, not medical, legal, or compliance advice: clinical accuracy and privacy rules stay with your organization.

Which patient education videos suit an AI avatar?

Videos that say the same thing to every patient:

  • Visit preparation, aftercare steps, and how-to explainers your clinical team has approved.
  • Front-desk and waiting-room information, such as check-in steps or the services a clinic offers.
  • Not a diagnosis, a test result, or anything about one patient. The avatar can't answer questions live either: each video is a job you submit, poll, and read.

How do I make a patient education video with an AI avatar?

Build a small library, one topic per video:

  • Create one presenter for the clinic from a prompt, a profile, or a photo, and reuse its avatar_handle so every video has the same face. How to create a reusable AI avatar shows the request.
  • Keep each script to one topic. Sume accepts an estimated 4–60 seconds, which today is roughly 165 words; how many words fit in 60 seconds explains the count.
  • Pick 16:9 for a waiting-room screen, or keep the default 9:16 for phones.
  • Create a preview with captions, let your reviewer approve the first-frame still, then call generate-video. Stored captions are burned in at that step; in current code the direct talking-video route refuses a captions field.
  • Put the topic and version in each Idempotency-Key. The docs say to reuse a key only for the same operation and payload, so an updated script gets a new key.
curl -X POST https://api.sume.com/v1/avatar-video-previews \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: clinic-check-in-v1" \
  -d '{
    "avatar_handle": "clinic_host",
    "script": "Welcome to our clinic. Please check in at the front desk and have your photo ID ready. We will call your name when your room is ready.",
    "scene": { "type": "prompt", "prompt": "Calm clinic reception, soft daylight" },
    "aspect_ratio": "16:9",
    "captions": { "enabled": true, "style": "slam" }
  }'

Can the videos be in Spanish or other languages?

Not from the same request: in current code, Avatar 1.0 talking videos speak English only. For another language, have your clinical team review the translated script, then speak it with TTS 1.0 and lip sync the same presenter with VEED Fabric 1.0, as Which languages can an AI avatar speak? shows. Fabric starts from the avatar's identity still, so that version won't carry the talking video's scene.

How do I keep patient information private?

Keep it out of the video. Sume returns avatar videos as public media.sume.com artifacts, so treat every file as readable by anyone who has the link:

  • Write scripts with no names, dates of birth, record numbers, or other details about a person.
  • Copy finished files to hosting your organization controls, such as a patient portal or the waiting-room player, and share that copy.
  • Keep API keys on your servers. Sume's docs say not to place them in frontend JavaScript or mobile apps, so a patient-facing page or app only gets the finished files.

How do I build a waiting-room loop?

Caption each clip first, join the clips into one file with Timeline 1.0, and set the screen's player to repeat it. Caption before the join, since today the caption job refuses a source over 60 seconds. In current code a render drops each clip's own sound, so use each clip's detached voice as the audio spine, as How to make an AI avatar video longer than 60 seconds shows.

What are the limits?

Avatar video bills per second by quality tier: $0.184/s standard, $0.245/s plus, $0.55/s max (no product image), plus a 5.5% agent fee by default. The other caps:

From Generate avatar video, Video captions, Timeline 1.0, and the Sume API reference, read 2026-09-28.
PieceLimit
One avatar videoEstimated 4–60 seconds, one avatar, one shared scene; English speech in current code
Shape and size1:1, 3:4, 9:16 (default), 4:3, or 16:9; 720p is the documented resolution
CaptionsInline captions up to an estimated 60 seconds; standalone jobs, today, need an audio stream and 60 seconds or less
Translated versionTTS up to 20,000 characters per request; Fabric takes 1–300 seconds of Sume-hosted audio, at most 10 MB
Waiting-room loopA 1–1,800 second spine, up to 200 slots and 20 audio parts, all Sume-hosted

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume