AI avatar for events: a host for screens and lobby loops

An AI avatar can host an event's scripted parts: the welcome, housekeeping, sponsor thanks, and session intros as short 16:9 clips, joined or looped.

5 min readSume
All posts

An AI avatar can host the scripted parts of an event: the welcome, housekeeping notes, sponsor thank-yous, session intros, and a lobby loop between sessions. On Sume, each part is a recorded talking video, not a live host. The avatar speaks the script you send, so anything unplanned, such as Q&A or a room change, still needs a person or a fresh render.

Sume facts come from the Generate avatar video and Timeline 1.0 docs, read on 2026-09-28; anything called current behavior is read from Sume's code. For a promo that sells tickets before the event, see how to make an event promo video with AI.

What can an AI event host present?

Anything you can script ahead of time, one segment per video:

  • A welcome and housekeeping: Wi-Fi, the day's schedule, exits, the code of conduct.
  • Sponsor thank-yous and session or speaker intros.
  • A lobby loop of the agenda, played between sessions.

How do I make the host's segments?

Render one POST /v1/avatar-1.0/talking-video job per segment:

  • Create the host once, or pick one from the stock avatar catalog, and send the same avatar_handle on every segment so every screen shows the same host.
  • Keep each script within an estimated 4–60 seconds, and name each Idempotency-Key after its slot in the run of show.
  • Set aspect_ratio to 16:9 for venue and stream screens. The default is 9:16.
  • Send the same scene on every segment: a stage prompt, or a public HTTPS photo of the real backdrop as a photo scene. Each video resolves one avatar and one shared scene.
  • Caption each finished segment for noisy rooms with POST /v1/video-captions. Today that job refuses a source over 60 seconds or one with no audio stream, so caption segments before you join them.
curl -X POST https://api.sume.com/v1/avatar-1.0/talking-video \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: summit-day1-welcome-v1" \
  -d '{
    "avatar_handle": "event_host",
    "script": "Good morning and welcome to day one. Wi-Fi details are on your badge, and the keynote starts at nine in Hall A.",
    "scene": { "type": "prompt", "prompt": "Conference stage, blue backdrop, soft spotlights" },
    "aspect_ratio": "16:9",
    "quality": "plus"
  }'

How do I join segments into a run of show or a lobby loop?

Join the captioned segments in order with one Timeline 1.0 render per file, as AI avatar news anchor does for a bulletin. Use each segment's detached voice as the audio spine, because in current code a render drops each clip's own sound. For screens, set output.width and output.height to a 16:9 size such as 1920×1080; Timeline renders 1080×1920 by default. For the lobby, render the loop once and set the screen's player to repeat the file.

Can an AI avatar host an event live?

No. Each segment is a job: you submit it, poll, and read a finished file, and a synchronous wait holds at most 30 seconds. The avatar can't hear the room, take questions, or react to a schedule change. Plan the live moments for a person, and render a new segment when the program changes. The host also speaks English only in current code; other languages go through TTS 1.0 and Fabric lip sync, as in which languages an AI avatar can speak.

What are the limits?

From Generate avatar video, Video captions, Timeline 1.0, and Sume's current code, read 2026-09-28.
PieceLimit
One segmentEstimated 4–60 seconds, one avatar, one shared scene; English speech in current code
Shape16:9 is one of five ratios (default 9:16); 720p is the documented resolution
CaptionsOne caption job per segment; today up to 60 seconds, with an audio stream
Joined fileA 1–1,800 second spine, 1–200 slots, up to 20 audio parts, all Sume-hosted

What does an AI event host cost?

Segments bill per second by quality tier: $0.184/s standard, $0.245/s plus, $0.55/s max (no product image). 5 minutes of segments at the default plus tier is $73.50 of avatar video, and each Timeline render is listed at $0.10 per output minute. Caption jobs and audio detach are priced per job; confirm them in GET /v1/catalog. All are plus a 5.5% agent fee by default.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume