AI avatar disclosure line as the first caption cue on a Sume clip
Tavus says disclosure features are still being worked on for realistic avatars. On a rendered Sume clip you can author the label yourself with caption cues.
To label an avatar clip as AI on Sume, burn an authored text cue into the video with standalone Video captions: send cues (or segments) with text, start and end, and no speech-to-text runs. The Tavus Griffin page says its model can deceive a person into believing it is not AI and that the company is working on safe disclosure features, so the label is something you should add yourself on any realistic avatar.
| Route | Source of text | Fit for a disclosure label |
|---|---|---|
| Inline captions on talking-video | Spoken script or video_inputs text | No: they transcribe what is said |
| Standalone video captions with script_text or STT | Speech | No, unless the label is spoken |
| Standalone video captions with cues or segments | Authored text with start and end | Yes |
How the cue route works
The docs describe cues (or segments) as authored overlay text, with text, start and end, burned without speech-to-text. That is exactly a fixed label. Use a Latin style for English text. If the label is Korean, select a Hangul style such as black-outline, because Sume rejects Korean text on slam, punch or tiktok-green with 400 caption_hangul_text_latin_style.
The standalone job takes a public HTTPS video_url, so run it on the finished avatar MP4. This is a second step after generation.
What the label should say
Keep it plain and early. Show it for the whole clip or at least the first three seconds, long enough to read at phone size.
- Say what it is: AI-generated presenter.
- Say who stands behind it if viewers could confuse it with a real person.
- Repeat it in the post text, where the platform allows.
What this does not settle
A caption cue is a visible label. It is not a legal opinion on the disclosure rules of any country or platform, and Sume's docs make no such claim. Check the rules for where the clip will run, and keep the label in the master file so every cut inherits it.
Order of operations
Generate the avatar clip first, with or without inline captions for speech. If you want both speech captions and a label, caption the speech inline, then run the standalone job on the resulting clean video with only the label cue, or send all cues together in one standalone job. Keep one master file that has the label so that every later edit inherits it.
Put the label's start at 0 and its end at the end of the clip for a short video. For a long one, a repeating label every 15 seconds costs nothing extra in the cue list, though it does take screen space.
Test the burned label on a small phone screen before publishing. Make sure it does not collide with platform interface elements at the bottom or top of the frame; the placement design fields can move it for styles that support design overrides.
Sources
Related posts
More in Use cases
- Roleplay team budget: Synthesia learner seats or a Sume clip library
Twelve-month cost of Synthesia Roleplay Training for 20 learners against a library of 24 scenario clips on Sume, and where each stops being the right tool.
- AI children's story video: 8 Omni scenes for $8.00
Eight 8-second Omni scenes at $1.00 each make a one-minute story on Sume. Use drawn characters: Google blocks minors in image edits in the EEA, UK and CH.
- AI dubbing workflow on Sume: detach, transcribe, TTS, timeline
Sume has no one-call dubbing endpoint. Chain detach, STT, your translation, TTS and a timeline render for a 60-second video at about $0.16 plus translation.
- AI presenter for YouTube Shorts: format limits and Sume output
Shorts can be square or vertical and run up to three minutes. Sume avatar videos are 9:16 or 1:1 at up to 60 seconds per job, so one job fits a Short.
Written by Sume