AI Avatar Museum Docent Clip: Script, Cost and Limits

A museum exhibit docent clip as an AI avatar render: a 3-scene script, the seconds math, request JSON and cost on Sume, plus why it is not a live guide.

4 min readSume
All posts

A museum label can only hold so many words. A short avatar docent clip, reached by QR code, can add the story behind a single object. This is a pre-rendered video on Sume Avatar 1.0, 4 to 60 seconds, not an interactive guide.

Prices are from Sume's API pricing page and the request shape from Generate avatar video, both read on 2026-10-05. The avatar must exist first: Create new avatar accepts a prompt, a profile or a photo.

What should the script and scenes look like?

One object per clip, three beats: what it is, why it matters, what to look for. Keep sentences short so each scene's 2.8 words per second estimate stays close to how long it takes to say.

The seconds column is Sume's planning estimate of ceil(words / 2.8) per scene, not a measured render time.

Scene plan and estimated seconds (read 2026-10-05)
SceneSpoken lineWordsSeconds
what"This bronze lamp was cast around the year eighte..."156
why"Lamps like this one lit shipyards at night, whic..."177
look"Look closely at the base, where a maker's mark i..."135
Total4518
POST https://api.sume.com/v1/avatar-1.0/talking-video
{
  "avatar_handle": "front_desk_host",
  "aspect_ratio": "9:16",
  "quality": "standard",
  "video_inputs": [
    { "id": "what", "voice": { "type": "text", "script": "This bronze lamp was cast around the year eighteen hundred, and it burned whale oil." }, "background": { "type": "prompt", "prompt": "Bright, plain room, locked-off camera, soft daylight" } },
    { "id": "why", "voice": { "type": "text", "script": "Lamps like this one lit shipyards at night, which is why this gallery sits beside the harbor." }, "background": { "type": "prompt", "prompt": "Bright, plain room, locked-off camera, soft daylight" } },
    { "id": "look", "voice": { "type": "text", "script": "Look closely at the base, where a maker's mark is worn almost smooth." }, "background": { "type": "prompt", "prompt": "Bright, plain room, locked-off camera, soft daylight" } }
  ]
}

What does it cost?

The plan above comes to 18 seconds, inside the 4 to 60 second window. At Sume's per-second rates (no product image), one render costs $3.31 on standard, $4.41 on plus (the default when quality is omitted) and $9.90 on max. Across 30 exhibit clips, that is $99.36, $132.30 and $297.00. Creating the avatar is a separate one-time $0.95; later clips reuse the handle.

Avatar Video, no product image, 18 seconds, 30 exhibit clips (rates read 2026-10-05)
QualityRate per secondOne clip30 exhibit clips
standard$0.184$3.31$99.36
plus$0.245$4.41$132.30
max$0.55$9.90$297.00

How does this compare with the live avatars in the news?

Tavus's Griffin page, read on 2026-10-05, describes Griffin-Lite as a research preview of a full-duplex model that is not available to customers. A visitor could one day talk to a docent; today's shipped path for a museum is a pre-rendered clip.

Because the same clip plays for every visitor, check every factual claim in the script against your curator's notes before rendering. The avatar reads the text; it does not verify it.

What will it not do?

It will not hear a visitor or answer a follow-up. Sume does not claim real-time conversation for Avatar 1.0, so questions still go to a human or a printed label.

Scripts for Avatar Video should be English; Sume's guidance is that non-English speech breaks caption alignment.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume