Hedra Avatar needs start frame and audio; Sume takes a handle and scri

Hedra Avatar generates from a start frame plus an audio track, up to 10 minutes. Sume's talking-video takes an avatar handle and a script, up to 60 seconds.

5 min readSume
All posts

Hedra Avatar needs two inputs, a start frame and an audio track, and its model page lists a maximum duration of 10 minutes. Sume's avatar talking-video needs a ready avatar_handle and a script, and accepts videos whose estimated length is 4 to 60 seconds. So Hedra expects you to bring the voice; Sume writes it for you.

Hedra's page is Hedra Avatar: AI Talking Avatar Video; Sume's are Generate avatar video and Create new avatar. Read 2026-10-02.

What does Hedra Avatar list?

The model page says it is text, image and audio to video, with the start frame and audio required. Aspect ratios are 1:1, 4:3, 3:4, 16:9, 9:16, 9:21 and 21:9, resolutions are 540p, 720p and 1080p, and the maximum duration is 10 minutes. Prices listed are 2.5 cents per second at 540p, 5 cents at 720p and 6.25 cents at 1080p. It is available through the Hedra Avatar API or the Creative Studio.

Note that a Hedra post on X describes the Avatar model as up to 5 minutes uncut, which disagrees with the model page. I used the model page; confirm the limit in the API before you plan long clips.

What does Sume need?

First you create a reusable avatar with POST /v1/avatar-1.0/generate, from a prompt, a profile (props) or a reference photo (photo with a public HTTPS image_url). The avatar is a job; when it finishes you get a handle. Then POST /v1/avatar-1.0/talking-video takes the handle and a script, with quality of standard, plus (default) or max, aspect_ratio of 1:1, 3:4, 9:16, 4:3 or 16:9, and resolution of 720p.

There is no start-frame field; the avatar is the starting identity. There is also no audio-file field on this route as documented.

Input comparison

The two designs split work differently: Hedra's job is lip-sync to your audio, Sume's is script to finished clip.

Avatar inputs, read 2026-10-02
ItemHedra AvatarSume Avatar 1.0
Required inputsStart frame and audioReady avatar handle and script or video_inputs
Max length10 minutes (model page)4-60 seconds estimated
Resolutions540p, 720p, 1080p720p
Aspect ratios1:1, 4:3, 3:4, 16:9, 9:16, 9:21, 21:91:1, 3:4, 9:16, 4:3, 16:9
VoiceYou supply the audioGenerated from the script

Which should I pick?

If you already have recorded narration, need more than a minute of one take, or need 1080p, Hedra's model page matches. If you start from text, want the voice and lip-sync in one request, and need many short clips with the same presenter, Sume's route is simpler: one avatar, many scripts, one idempotent request each. For a longer Sume piece, read AI avatar video longer than 60 seconds and plan to join several clips.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume