Wedding Venue Tour Intro With an AI Avatar: Script and Cost
Build a 9:16 avatar intro for a wedding venue inquiry reply: scene script, seconds per scene, request JSON, and the cost of each Sume quality tier.
A venue gets the same first questions from every couple: capacity, dates, what is included. A short avatar host video can answer them in the reply email. On Sume Avatar 1.0 that is a scripted clip of up to 60 seconds with one avatar and one shared scene.
Prices are from Sume's API pricing page and the request shape from Generate avatar video, both read on 2026-10-05. The avatar must exist first: Create new avatar accepts a prompt, a profile or a photo.
What should the script and scenes look like?
Because every scene must resolve to one shared look, write the scenes as beats in the same room. Each line is counted at 2.8 words per second, rounded up per scene.
The seconds column is Sume's planning estimate of ceil(words / 2.8) per scene, not a measured render time.
| Scene | Spoken line | Words | Seconds |
|---|---|---|---|
| hello | "Thanks for asking about Alder Hall. Here's the q..." | 16 | 6 |
| space | "The ballroom seats one hundred and twenty for di..." | 22 | 8 |
| next | "Reply with two dates you like, and we'll hold th..." | 16 | 6 |
| Total | 54 | 20 |
POST https://api.sume.com/v1/avatar-1.0/talking-video
{
"avatar_handle": "front_desk_host",
"aspect_ratio": "16:9",
"quality": "plus",
"video_inputs": [
{ "id": "hello", "voice": { "type": "text", "script": "Thanks for asking about Alder Hall. Here's the quick version of what couples ask us first." }, "background": { "type": "prompt", "prompt": "Bright, plain room, locked-off camera, soft daylight" } },
{ "id": "space", "voice": { "type": "text", "script": "The ballroom seats one hundred and twenty for dinner, and the garden holds a ceremony for up to one hundred and fifty." }, "background": { "type": "prompt", "prompt": "Bright, plain room, locked-off camera, soft daylight" } },
{ "id": "next", "voice": { "type": "text", "script": "Reply with two dates you like, and we'll hold them for seven days while you visit." }, "background": { "type": "prompt", "prompt": "Bright, plain room, locked-off camera, soft daylight" } }
]
}What does it cost?
The plan above comes to 20 seconds, inside the 4 to 60 second window. At Sume's per-second rates (no product image), one render costs $3.68 on standard, $4.90 on plus (the default when quality is omitted) and $11.00 on max. Across 25 inquiry replies, that is $92.00, $122.50 and $275.00. Creating the avatar is a separate one-time $0.95; later clips reuse the handle.
| Quality | Rate per second | One clip | 25 inquiry replies |
|---|---|---|---|
| standard | $0.184 | $3.68 | $92.00 |
| plus | $0.245 | $4.90 | $122.50 |
| max | $0.55 | $11.00 | $275.00 |
Should the avatar be on location?
Not necessarily. A prompt-based background gives the avatar a plausible room, but it is not your real ballroom. If couples need to see the actual space, put real photos beside the clip in the email, or use a photo scene as described in Generate avatar video.
Say the capacity numbers in the script only if they are current; the avatar will read whatever you write.
What will it not do?
It will not walk through a real venue or react to a couple's questions. Sume Avatar 1.0 renders what the script says and nothing more, so live Q&A stays with your coordinator.
Scripts for Avatar Video should be English; Sume's guidance is that non-English speech breaks caption alignment.
Sources
Related posts
More in Use cases
- What is a critical error in TTS? A rubric after Nova 2 Sonic's 28%
Amazon reports 28% fewer critical errors in Nova 2 Sonic on an internal set and does not define them. Define yours, then tally with a read-back on Sume.
- What counts as original for an AI-made YouTube Short?
YouTube's Oct 2026 Shorts update favors original work. Here is what its own pages say about AI, templates and edits, and what they leave unsaid.
- What size should a TikTok video be? 2026 ad minimums by type
TikTok's own ad pages list 9:16 at 540x960 minimum for in-feed and 720x1280 for App Bundle. See the table and the Sume parameters that hit each one.
- Which AI video model for ads, product shots or talking heads?
On Sume, pick Wan 3.0 or Omni for ads, Kling 3 for silent product shots, and the Avatar Video route for talking heads. One table with rates and limits.
Written by Sume