AI clone from 2 minutes of video or one photo: what each needs
Tavus asks for two minutes of 1080p video and written consent. Sume Avatar 1.0 starts from a prompt, traits or one photo URL. The inputs compared, dated.

Tavus needs about two minutes of video, including at least 30 seconds of speech. Sume Avatar 1.0 needs one public HTTPS photo, a text prompt, or a few profile traits. They produce different things: a Tavus replica is built to hold live conversations, while a Sume avatar is built to speak a script you send, in a finished video.
The Tavus numbers come from its October 2 guide, How to create an AI clone from 2 minutes of video, read 2026-10-06. The Sume numbers come from the avatar docs and the pricing code on main.
Input requirements, side by side
| Item | Tavus AI clone | Sume Avatar 1.0 |
|---|---|---|
| Source material | About two minutes of video, at least 30 seconds of speech | A prompt, props (ethnicity, sex, age) or one image_url |
| Video quality | 1080p at 25 fps or better, recorded in a desktop app | No video: a fetchable public HTTPS image |
| Still segment | At least 30 seconds of silent, still footage | None |
| Camera and look | Eye level, face filling at least a quarter of the frame, soft light; no glasses, jewelry or patterned clothing | Image must be fetchable; Sume rejects localhost, private-network, non-HTTPS and non-image URLs before it submits |
| Training time | Not stated; points to its training docs | Creation is a job; poll it until it completes |
| Consent | Documented consent required before cloning | Your responsibility: use only photos you have the right to use |
| Cost | Not given in the article | $0.95 once per avatar |
What changes in practice
A clone recorded at 1080p carries listening posture and person-specific voice detail, which suits a face that talks back. If what you need is a presenter who reads weekly updates, recording and consent logistics are the heavy part, and a photo-based avatar removes most of them.
The trade is control of the performance. With Sume you steer a video through its script, an optional scene prompt or scene photo, a quality tier and, if wanted, inline captions. Resolution is 720p at this time, and a video runs 4-60 seconds.
Consent still applies to a photo
Tavus says explicit, documented consent is mandatory before cloning, and notes EU Article 50 disclosure for AI-generated likenesses since 2 August 2026. That reasoning does not depend on how the likeness is made. If the photo shows a real person, get their written permission before you create the avatar, and keep the record next to the avatar handle.
Sources
Related posts
More in Sume Avatar 1.0
- Digital twin vs AI replica vs AI avatar: which word is right?
Tavus says an AI copy of a person is not a digital twin. Four system types in one table, and where a Sume Avatar 1.0 handle sits among them.
- Face swap, motion control or lip sync: which Sume avatar route?
Pick the Sume avatar route by what you already have: a clip with a person (face swap), a motion video (Kling motion control) or a voice line (H3 Max lip sync).
- Interview video with one AI avatar host and guest cards on screen
A final avatar video holds one avatar, so an interview needs a host clip plus guest stills. Use Timeline compose at $0.02 a shot to put a still beside the host.
- Pipecat avatar TTFB metrics vs timing a Sume avatar job
Pipecat's Tavus service reports TTFB from TTSStartedFrame and BotStartedSpeakingFrame. A Sume avatar job has no first byte to time: measure submit to completed.
Written by Sume