Corporate talking head video with AI: CEO and team updates
Make a corporate talking head video with AI: script the message, pick a presenter, render it in 16:9 or 9:16, and host the file where staff sign in.

A corporate talking head video is a short clip of one presenter speaking to camera for a company audience: a CEO message, a policy change, or a quarterly update for employees. With AI, you write the script and render a presenter speaking it, so nobody has to be filmed. The presenter can be a stock or generated person, a likeness generated from an executive's photo with their consent, or the executive's own photo lip-synced to a clone of their voice.
Sume facts come from the Generate avatar video and Create new avatar docs and the Sume API reference, read on 2026-09-28. Anything called current behavior is read from Sume's code. Using a real person's face or voice needs their consent, and how you label AI video is your company's decision.
Which presenter should a company video use?
Choose by whose message it is. All four routes end in a talking video; they differ in whose face and voice it carries. The last one is the only route in your executive's own voice, and How to clone yourself with AI walks through it.
| Presenter | Sume route | What to know |
|---|---|---|
| A stock presenter | POST /v1/avatar-catalog/search, then the chosen avatar's handle | Searches public system avatars plus your workspace's own; nothing to create |
| A generated presenter | POST /v1/avatar-1.0/generate with a prompt or a props profile | A reusable avatar under a handle you choose |
| An executive's likeness | The same route with a photo input | Current code redraws the photo as a new portrait and generates a voice to suit the face, so it isn't their voice |
| The executive's own voice and photo | Voice cloned in the Sume app, POST /v1/tts-1.0/generate, then POST /v1/veed/fabric-1.0 | Fabric animates the photo you send to 1–300 seconds of Sume-hosted audio |
How do I make a CEO video message with AI?
Treat each message as one render job:
- Write the script. Sume accepts a script it estimates at 4–60 seconds, which today is roughly 165 words; how many words fit in 60 seconds explains the count.
- Create the presenter once and send the same
avatar_handleevery time, so each update has the same face. - Set
aspect_ratioto16:9for an intranet page or a town-hall screen. The default,9:16, suits phones. - Set the room with
scene: a text prompt, or a public HTTPS photo of your office as aphotoscene. - To approve the framing before paying for the full render, start from a first-frame preview.
curl -X POST https://api.sume.com/v1/avatar-1.0/talking-video \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: q3-update-ceo-v1" \
-d '{
"avatar_handle": "company_host",
"script": "Hi everyone. Our third quarter closed this week, and I want to thank every team for the work behind it.",
"scene": { "type": "prompt", "prompt": "Modern office, bookshelf behind, soft window light" },
"aspect_ratio": "16:9",
"quality": "plus"
}'How do I send a message longer than a minute?
Split it into parts of up to 60 seconds, render each with the same handle, scene, and ratio, and join them with Timeline 1.0 as How to make an AI avatar video longer than 60 seconds shows, with each part's detached voice as the audio spine: in current code a render drops each clip's own sound.
Who can see the finished video?
Completed avatar videos come back as public media.sume.com artifacts, so treat each URL as open to anyone who has it. For a message meant only for staff:
- Download the MP4 and publish it where employees sign in, such as your intranet or internal video platform.
- Keep the
media.sume.comlink out of public channels, and read Do AI-generated video URLs expire? before you rely on it.
What are the limits?
- English speech only, in current code: the clip prompt asks for English, and avatar voices are cloned in English. Other languages go through TTS 1.0 with
language, then Fabric lip sync, as in which languages an AI avatar can speak. - One presenter and one shared scene per video, so a two-person conversation is one render per turn, cut together, as in AI avatar conversation videos.
- 720p is the documented resolution, in
1:1,3:4,9:16,4:3, or16:9. - It's a recorded file, not a live stream: each video is a job you submit, poll, and read.
What does a corporate avatar video cost?
Avatar video bills per second by quality tier: $0.184/s standard, $0.245/s plus, $0.55/s max (no product image). A 60-second message at the default plus tier is $14.70 of avatar video, and creating an avatar costs $0.95 per avatar. The own-voice route bills TTS at $0.0475 per 1,000 characters and Fabric at $0.1875 per audio second (720p) instead. Each is plus a 5.5% agent fee by default.
Sources
Related posts
More in Use cases
- Creatify alternatives: link, image or script to video ad
Creatify alternatives compared on what they start from (a product link, product images, or a script) and whether batches run over an API. Dated table.
- Customer onboarding video: one short clip per setup step
A customer onboarding video walks a new customer through one step of getting started. Render one short AI avatar clip per step from your backend.
- Direct response video ads: what they are and how to make them
A direct response video ad asks for one action now, such as buy or sign up, and is judged by that response. Its structure, and how to make variants with AI.
- Does TikTok allow AI-generated content? Yes, with a label
Yes. TikTok allows AI-generated content but requires a label on realistic AI images, audio or video. Via the API, is_aigc: true adds the label.
Written by Sume