Vacation rental check-in video: an avatar host with silence beats
One avatar, one shared background and silence beats let a host walk guests through check-in in 60 seconds or less. Set video_inputs on Avatar Video.
To make a vacation rental check-in video with an avatar host, send ordered video_inputs to POST /v1/avatar-1.0/talking-video: spoken scenes for the steps, voice.type: "silence" beats where a guest would look at the lockbox or the thermostat, and the same background in every scene. Total length must land in the 4 to 60 second window.
Current execution supports one avatar per final video and expects every scene background to resolve to one shared scene. Source: Avatar videos, read 2026-10-04.
Scene plan
Each scene has an id, a voice and a background. A spoken scene is type: "text" with exactly one of script or input_text and a duration. A silence beat needs duration and no script.
| Scene id | Voice | Seconds |
|---|---|---|
| welcome | text | 6 |
| lockbox | text | 8 |
| pause-lockbox | silence | 3 |
| wifi | text | 7 |
| checkout | text | 8 |
| Total | 32 |
Request
The total, 6 + 8 + 3 + 7 + 8, is 32 seconds, inside the 4 to 60 window. Set aspect_ratio to 16:9 for a listing page, since the default is 9:16. Choose the standard tier for a draft.
curl -X POST https://api.sume.com/v1/avatar-1.0/talking-video \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: checkin-video-001" \
-d '{
"avatar_handle": "host_maria",
"aspect_ratio": "16:9",
"quality": "standard",
"video_inputs": [
{ "id": "welcome", "voice": { "type": "text", "script": "Welcome! I am Maria. Here is how check-in works.", "duration": 6 },
"background": { "type": "prompt", "prompt": "Bright entryway of a beach cottage" } },
{ "id": "lockbox", "voice": { "type": "text", "script": "The lockbox is by the front door. Your code is in your booking message.", "duration": 8 },
"background": { "type": "prompt", "prompt": "Bright entryway of a beach cottage" } },
{ "id": "pause-lockbox", "voice": { "type": "silence", "duration": 3 },
"background": { "type": "prompt", "prompt": "Bright entryway of a beach cottage" } },
{ "id": "wifi", "voice": { "type": "text", "script": "Wi-Fi details are on the fridge.", "duration": 7 },
"background": { "type": "prompt", "prompt": "Bright entryway of a beach cottage" } },
{ "id": "checkout", "voice": { "type": "text", "script": "At checkout, lock the door and leave the key in the box. Enjoy your stay.", "duration": 8 },
"background": { "type": "prompt", "prompt": "Bright entryway of a beach cottage" } }
]
}\Keep secrets out
An avatar host is not a person who will meet the guest, so say so in the listing text. Keep the codes and addresses out of the video itself, because the video lives longer than the stay.
Preview and captions
To approve the first frame before paying for a render, use Avatar video previews with the same body, and call generate-video when the still looks right. Inline captions on talking-video are not billed separately and apply to videos of up to 60 seconds.
Sources
Related posts
More in Use cases
- Vendor pages to read before AI music goes in an ad
The vendor pages that answer commercial-use questions for Suno v6, ElevenLabs Music v2.5 and Google Lyria 3.5, and which question each answers, read 2026-10-04.
- Vietnam AI Law Article 11: labelling real people, real events
Under Vietnam's AI Law, deployers must label AI media that imitates a real person's face or voice or recreates a real event. Art and film get a softer rule.
- Vietnam AI Law Article 11: machine-readable marks for AI media
Vietnam's Law No. 134/2025/QH15, in force 1 March 2026, makes AI providers mark audio, image and video in machine-readable form. Sume docs show no such mark.
- Voicemail drop: pick immediate or delayed detection
ElevenLabs added voicemail detection for Twilio calls with immediate and delayed modes. Record the message ahead of time with Sume's tts_create and drop it.
Written by Sume