Captions AI Twin from a video upload vs Sume's photo avatar
Captions can build an AI Twin from an uploaded video. Sume's avatar API takes a prompt, a profile or a public HTTPS reference image, but no video input.
Sume cannot create an avatar from an uploaded video: the avatar API accepts a prompt, a structured profile, or one reference image. Captions, per its release notes, lets you set up an AI Twin by uploading a video, and later from a single still photo.
Captions facts are from its release notes (an undated list); Sume facts from Create new avatar, read 2026-10-01.
What does the Captions app accept?
Two entries: AI Twin creation from a video upload, not just a recording in the app, and AI Twin from a single still photo with no recording required.
What does Sume accept?
Three input kinds on POST /v1/avatar-1.0/generate: Prompt (describe the avatar), Profile (structured traits, sent as props), and Image (a reference image, sent as photo). The image_url must be a fetchable public HTTPS image URL; localhost, private-network, non-HTTPS and non-image responses are rejected before submission.
| Source | Captions app | Sume avatar API |
|---|---|---|
| Video upload | Yes (release notes) | Not an input type |
| Single photo | Yes (release notes) | photo with image_url |
| Text only | Not in the notes | prompt |
How do I create an avatar from a photo?
Send avatar_handle and input: { "type": "photo", "image_url": "https://..." } with an Idempotency-Key. Each request creates a job; poll it, then use the returned handle for avatar videos. If you only have a video, export a clean still frame of the face and host that image. Compare Mirage's AI Twin.
Sources
Related posts
More in Use cases
- Captions Avatar Looks in Prompt to Video: the Sume handle workflow
Captions saves looks inside one avatar for Prompt to Video. For several looks in Sume, make one avatar_handle per look and render one job per final video.
- Captions horizontal 16:9 AI Edit vs Sume avatar aspect ratios
Captions AI Edit now outputs 16:9. Sume avatar videos accept 16:9 too, at 720p only, with a script or plan estimated at 4-60 seconds. 9:16 is the default.
- Captions app SRT export vs Sume burned-in caption cues
The Captions app can export an SRT file for external caption tracks. Sume burns captions into the MP4 from authored cues with text, start and end.
- Captions AI avatar looks vs one Sume avatar handle per look
Captions added Avatar Looks to save new looks per avatar. On Sume you create one avatar handle per look from a prompt, a profile or a reference image.
Written by Sume