Captions AI Twin from a video upload vs Sume's photo avatar

Captions can build an AI Twin from an uploaded video. Sume's avatar API takes a prompt, a profile or a public HTTPS reference image, but no video input.

4 min readSume
All posts

Sume cannot create an avatar from an uploaded video: the avatar API accepts a prompt, a structured profile, or one reference image. Captions, per its release notes, lets you set up an AI Twin by uploading a video, and later from a single still photo.

Captions facts are from its release notes (an undated list); Sume facts from Create new avatar, read 2026-10-01.

What does the Captions app accept?

Two entries: AI Twin creation from a video upload, not just a recording in the app, and AI Twin from a single still photo with no recording required.

What does Sume accept?

Three input kinds on POST /v1/avatar-1.0/generate: Prompt (describe the avatar), Profile (structured traits, sent as props), and Image (a reference image, sent as photo). The image_url must be a fetchable public HTTPS image URL; localhost, private-network, non-HTTPS and non-image responses are rejected before submission.

Source media for a digital twin, read 2026-10-01.
SourceCaptions appSume avatar API
Video uploadYes (release notes)Not an input type
Single photoYes (release notes)photo with image_url
Text onlyNot in the notesprompt

How do I create an avatar from a photo?

Send avatar_handle and input: { "type": "photo", "image_url": "https://..." } with an Idempotency-Key. Each request creates a job; poll it, then use the returned handle for avatar videos. If you only have a video, export a clean still frame of the face and host that image. Compare Mirage's AI Twin.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume