AI clone from 2 minutes of video or one photo: what each needs

Tavus asks for two minutes of 1080p video and written consent. Sume Avatar 1.0 starts from a prompt, traits or one photo URL. The inputs compared, dated.

5 min readSume
All posts

Tavus needs about two minutes of video, including at least 30 seconds of speech. Sume Avatar 1.0 needs one public HTTPS photo, a text prompt, or a few profile traits. They produce different things: a Tavus replica is built to hold live conversations, while a Sume avatar is built to speak a script you send, in a finished video.

The Tavus numbers come from its October 2 guide, How to create an AI clone from 2 minutes of video, read 2026-10-06. The Sume numbers come from the avatar docs and the pricing code on main.

Input requirements, side by side

Tavus recording rules (read 2026-10-06) against Sume Avatar 1.0 inputs from the docs on main.
ItemTavus AI cloneSume Avatar 1.0
Source materialAbout two minutes of video, at least 30 seconds of speechA prompt, props (ethnicity, sex, age) or one image_url
Video quality1080p at 25 fps or better, recorded in a desktop appNo video: a fetchable public HTTPS image
Still segmentAt least 30 seconds of silent, still footageNone
Camera and lookEye level, face filling at least a quarter of the frame, soft light; no glasses, jewelry or patterned clothingImage must be fetchable; Sume rejects localhost, private-network, non-HTTPS and non-image URLs before it submits
Training timeNot stated; points to its training docsCreation is a job; poll it until it completes
ConsentDocumented consent required before cloningYour responsibility: use only photos you have the right to use
CostNot given in the article$0.95 once per avatar

What changes in practice

A clone recorded at 1080p carries listening posture and person-specific voice detail, which suits a face that talks back. If what you need is a presenter who reads weekly updates, recording and consent logistics are the heavy part, and a photo-based avatar removes most of them.

The trade is control of the performance. With Sume you steer a video through its script, an optional scene prompt or scene photo, a quality tier and, if wanted, inline captions. Resolution is 720p at this time, and a video runs 4-60 seconds.

Consent still applies to a photo

Tavus says explicit, documented consent is mandatory before cloning, and notes EU Article 50 disclosure for AI-generated likenesses since 2 August 2026. That reasoning does not depend on how the likeness is made. If the photo shows a real person, get their written permission before you create the avatar, and keep the record next to the avatar handle.

Sources

Related posts

More in Sume Avatar 1.0

All Sume Avatar 1.0 posts

Written by Sume