AI avatar prompts: how to describe a talking presenter
An AI avatar prompt for video describes one person: age, look, hair, clothing, and expression. What to write, what Sume adds, and what it can't set.
An AI avatar prompt for a talking presenter describes one person in plain words: rough age, look, hair, clothing, and expression. On Sume, that text is the prompt input of POST /v1/avatar-1.0/generate, and in current code Sume adds its own direction for a realistic vertical smartphone portrait with a clean, simple background, so cartoon styles, text, and logos work against it. The prompt sets the face and the outfit; the voice and each video's setting come later.
Facts come from Sume's Create new avatar docs and the Sume API reference, read on 2026-09-28. What Sume adds to a prompt is current code, not a documented contract, and no prompt guarantees a look. Looking for profile-picture prompts for an image tool instead? See AI headshot generator.
What should an AI avatar prompt include?
One person, described the way a casting note would describe them:
- Age and look: "a woman in her early thirties".
- Hair: "shoulder-length curly dark hair".
- Clothing: "a navy blazer over a plain white T-shirt".
- Expression: "a warm, relaxed smile".
- Together: "A woman in her early thirties with shoulder-length curly dark hair, wearing a navy blazer over a plain white T-shirt, with a warm, relaxed smile."
What does Sume add to my prompt?
In current code, Sume appends a fixed direction to every prompt avatar: a realistic vertical smartphone portrait, natural daylight, an approachable expression, and a clean, simple background. It also asks the image model to avoid text, logos, extra people, heavy filters, and stylized illustration, and it renders the portrait at 9:16. So:
- Skip style words such as anime, cartoon, or oil painting; they pull against the realism direction.
- Leave out rooms and props. The portrait background is meant to stay plain, and each video's setting comes from its own
scene, as in UGC video prompt for AI avatars. - Describe one person in plain clothing, without printed slogans or brand marks.
What can't an avatar prompt control?
- The voice. In current code it is generated afterwards to suit the person's look and age, then cloned in English; the prompt has no voice setting.
- Gestures, tone, and camera moves: talking videos have no fields for them.
- Later changes. The public avatar routes only list and read avatars, and reusing a handle answers
409 avatar_handle_taken, so a new look needs a new handle.
Should I use a prompt, a profile, or a photo?
Pick by how much you need to control. If you only need an age, sex, and ethnicity mix, the profile input sends exactly those three values, and current code writes the portrait description from them for you.
| Input | `input.type` | What you send | Use it when |
|---|---|---|---|
| Prompt | prompt | A text description of the person | You want to set hair, clothing, and expression |
| Profile | props | ethnicity (one of 8 values), sex (male or female), age (20–80) | A demographic mix is enough; current code frames it as an upper-body presenter portrait |
| Photo | photo | A public HTTPS image_url | The avatar should resemble someone who agreed to it; current code redraws the photo as a new portrait |
How do I send the prompt, and what does it cost?
Send it with a handle of your choice. The job returns a reusable avatar; poll it as How to create a reusable AI avatar shows.
avatar_handle: 2–30 letters, digits, underscores, or periods, with no hyphens and no period or underscore first, last, or twice in a row. Thesume_prefix is reserved for Sume's own avatars.- When the avatar is ready, look at its
preview_image_urlbefore you render any videos. Itsvoice.statusmust bereadyas well; in current code a talking video requested earlier answers409 avatar_not_ready. - Creating an avatar costs $0.95 per avatar, plus a 5.5% agent fee by default.
curl -X POST https://api.sume.com/v1/avatar-1.0/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: avatar-prompt-host-001" \
-d '{
"avatar_handle": "brand_host",
"input": {
"type": "prompt",
"prompt": "A woman in her early thirties with shoulder-length curly dark hair, wearing a navy blazer over a plain white T-shirt, with a warm, relaxed smile"
}
}'Sources
Related posts
More in Sume Avatar 1.0
- Kling API avatar: one image plus audio, 2 to 300 seconds
Kling's API has an Avatar endpoint: one reference image plus a 2–300 second audio track becomes a talking video, billed per second. Inputs and Sume options.
- Difference between lip sync and dubbing: which do you need?
Dubbing replaces the speech in a video; lip sync matches a mouth to the audio. Lip-sync dubbing does both. What each means, and which one you need.
- How long should a microlearning video be? Length and AI
No fixed standard: one objective per video, and research on course videos favors 6 minutes or less. How to size lessons and make them with an AI avatar.
- Real-time AI avatar vs video avatar API: the difference
A real-time AI avatar talks live in a session; a video avatar API renders a finished file from a script. Which vendors sell which, and where Sume fits.
Written by Sume