AI avatar for a customer service kiosk: clips, not live chat

A kiosk that answers any question needs a live avatar. A kiosk with known answers can play rendered clips. What Sume's avatar video can and cannot do there.

4 min readSume
All posts

Sume can supply the video for a customer service kiosk only if the answers are known in advance. Its avatar video is a job that turns a script into a clip of 4 to 60 seconds; nothing in the docs describes a live session, so it cannot answer a question nobody scripted. A kiosk that must reply to anything a visitor says needs a live avatar such as Google's, which streams video on a live dialogue model.

So the design question is how many of your visitors' questions you can list. If it is a dozen, clips fit. If it is open-ended, they do not.

Which kiosk jobs fit rendered clips?

Kiosk jobs and clip fit, read 2026-09-29.
Kiosk jobRendered clip?
Welcome greetingYes, one script
Store hours, directions, returns policyYes, one clip per known question
Step-by-step how-toYes, video_inputs scenes up to 60 seconds
Answer to an unscripted questionNo, needs a live dialogue model

How do I build the clip set?

Write one script per answer and send each as its own POST /v1/avatar-1.0/talking-video request with the same avatar_handle, so the presenter looks the same on every screen. Use quality of standard, the tier the docs describe as the quickest Sume path, while you draft, then max for the clips you keep.

Set aspect_ratio to match the screen: 9:16 is the default, and 16:9, 4:3, 3:4 and 1:1 are also accepted. Output is currently 720p.

How long should each clip be?

Each job must land in the 4 to 60 second window, and a kiosk answer rarely needs more than the low end of it. Short clips also keep the visitor's wait short and make one wrong line cheap to re-render.

For a longer walkthrough, use ordered video_inputs with one scene per step. A scene with voice.type of silence and a duration gives the visitor a beat to look at the screen. Total planned duration still has to fit the window; see multi-scene avatar video.

How do I get the files onto the kiosk?

Each request returns a job. Poll it, or send a webhook; when it is complete the result can include public media.sume.com video artifacts. Download the MP4s and bundle them with your kiosk software. The docs do not describe a player or embed, so playback is yours to build.

What should the kiosk do when the clip set has no answer?

Show a fallback: a text prompt, a call button or a hand-off to a person. A recorded avatar cannot improvise, and making it appear to would be misleading. Keep the fallback visible, and add a new clip when the same unanswered question keeps coming up.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume