Hedra Avatar needs start frame and audio; Sume takes a handle and scri
Hedra Avatar generates from a start frame plus an audio track, up to 10 minutes. Sume's talking-video takes an avatar handle and a script, up to 60 seconds.
Hedra Avatar needs two inputs, a start frame and an audio track, and its model page lists a maximum duration of 10 minutes. Sume's avatar talking-video needs a ready avatar_handle and a script, and accepts videos whose estimated length is 4 to 60 seconds. So Hedra expects you to bring the voice; Sume writes it for you.
Hedra's page is Hedra Avatar: AI Talking Avatar Video; Sume's are Generate avatar video and Create new avatar. Read 2026-10-02.
What does Hedra Avatar list?
The model page says it is text, image and audio to video, with the start frame and audio required. Aspect ratios are 1:1, 4:3, 3:4, 16:9, 9:16, 9:21 and 21:9, resolutions are 540p, 720p and 1080p, and the maximum duration is 10 minutes. Prices listed are 2.5 cents per second at 540p, 5 cents at 720p and 6.25 cents at 1080p. It is available through the Hedra Avatar API or the Creative Studio.
Note that a Hedra post on X describes the Avatar model as up to 5 minutes uncut, which disagrees with the model page. I used the model page; confirm the limit in the API before you plan long clips.
What does Sume need?
First you create a reusable avatar with POST /v1/avatar-1.0/generate, from a prompt, a profile (props) or a reference photo (photo with a public HTTPS image_url). The avatar is a job; when it finishes you get a handle. Then POST /v1/avatar-1.0/talking-video takes the handle and a script, with quality of standard, plus (default) or max, aspect_ratio of 1:1, 3:4, 9:16, 4:3 or 16:9, and resolution of 720p.
There is no start-frame field; the avatar is the starting identity. There is also no audio-file field on this route as documented.
Input comparison
The two designs split work differently: Hedra's job is lip-sync to your audio, Sume's is script to finished clip.
| Item | Hedra Avatar | Sume Avatar 1.0 |
|---|---|---|
| Required inputs | Start frame and audio | Ready avatar handle and script or video_inputs |
| Max length | 10 minutes (model page) | 4-60 seconds estimated |
| Resolutions | 540p, 720p, 1080p | 720p |
| Aspect ratios | 1:1, 4:3, 3:4, 16:9, 9:16, 9:21, 21:9 | 1:1, 3:4, 9:16, 4:3, 16:9 |
| Voice | You supply the audio | Generated from the script |
Which should I pick?
If you already have recorded narration, need more than a minute of one take, or need 1080p, Hedra's model page matches. If you start from text, want the voice and lip-sync in one request, and need many short clips with the same presenter, Sume's route is simpler: one avatar, many scripts, one idempotent request each. For a longer Sume piece, read AI avatar video longer than 60 seconds and plan to join several clips.
Sources
Related posts
More in Comparisons
- Hedra API gives 75+ models in one place vs Sume Avatar 1.0 route
Hedra's developer page advertises 75+ models from 13 providers behind api.hedra.com. Sume Avatar 1.0 offers a narrow route. When breadth wins.
- Hedra's developer platform: API, SDK, CLI and MCP vs Sume
Hedra opened its models through an API, SDKs, a CLI and MCP on August 4, 2026. A map of what each surface covers and what Sume offers for avatar work.
- Hedra job status progress and estimated_completion_at vs Sume events
Hedra's v3 status endpoint returns progress and estimated_completion_at; Sume's job status gives queued, processing, completed plus an events timeline.
- HeyGen break tag: 5-second pause max, and a longer pause on Sume
HeyGen's professional voice clones accept a break tag with a 5-second cap per pause. On Sume an avatar video pause is a silence scene, up to 60 seconds.
Written by Sume