Mirage Avatar X: a 10-second AI twin vs Sume avatar inputs
Captions' Mirage Avatar X powers its avatars and AI twins, and a twin needs ten seconds of footage. Sume avatars come from a prompt, a profile or an image.
Mirage Avatar X is Captions' avatar model, announced July 28, 2026; it powers all Captions avatars and AI twins, and a twin takes ten seconds of source footage recorded in Captions. Sume creates avatars differently: from a prompt, a structured profile, or a reference image. A video clip is not one of the documented inputs.
What does Captions say about Avatar X?
The announcement says Avatar X generates the whole performance at once, stitching video, audio and motion as one scene, and that it supports horizontal or vertical video. It says a personal twin only takes ten seconds of source footage, and that you can instead pick a library avatar or generate custom avatars from prompts. These are vendor statements, read 2026-09-30.
How do you make an avatar on Sume?
The avatar docs list three ways: a prompt that describes the avatar, a profile of structured traits, or an image used as a reference. Each request creates a job; when it completes you use the returned avatar handle or resource id to make avatar videos. The handle may include a leading @, and Sume stores it normalized without it.
| Input | Captions (per its post) | Sume (per its docs) |
|---|---|---|
| Text prompt | Custom avatars from prompts | Prompt input |
| Structured traits | Not mentioned | Profile input |
| Photo | Not mentioned | Image input (reference image) |
| Video footage | Ten seconds for an AI twin | Not a documented input |
Can I turn my own footage into a Sume avatar?
Not as a video input. If you have a clear still of the person, the image input is the documented route; see create an AI avatar from video for what to do with footage you already have. Only use likenesses you have the right to use.
What quality choices exist once I have a handle?
Avatar Video accepts quality of standard, plus or max. plus is the default, standard is the fastest Sume execution path, and max is the highest tier with slower turnaround. Details are in Avatar videos.
Sources
Related posts
More in Models
- Nano Banana 14 reference images vs Sume's 10 input_references
Google says Gemini 3 image models mix up to 14 reference images. Sume's Nano Banana models take up to 10 input_references; ChatGPT Image 2.5 takes 16.
- Nano Banana Google Search grounding: not a Sume parameter
Google documents Search grounding for its Gemini image models. Sume lists Nano Banana 2 but no grounding field, and unlisted fields return 400.
- Nano Banana Pro aspect ratio: auto keeps the reference shape
Freepik's changelog says Nano Banana Pro accepts aspect_ratio auto to follow reference images; omitting it stays 1:1. On Sume, omitting is also not auto.
- Nano Banana thinking mode and thought images on Sume
Google's Gemini 3 image models use thinking mode and uncharged thought images. Sume lists no thinking field, and you pay only for completed images.
Written by Sume