Hedra legacy avatar API needs 4 calls; Sume avatar video needs 1 POST
Hedra's legacy avatar guide creates and uploads two assets, then generates and polls. Sume's talking-video route is one POST with a handle and a script.
Hedra's legacy avatar guide has you upload audio, upload a portrait, start a generation with both asset ids, and poll until it completes. Sume's route is a single POST /v1/avatar-1.0/talking-video with an avatar_handle and a script, then polling the returned job.
The difference is where the inputs live. Hedra treats audio and image as files you place on its side first. Sume's request carries a reference to an avatar you created earlier and the text to speak.
The Hedra flow as documented
From the Hedra legacy avatar guide, read 2026-10-02:
- The guide lists
model_slugastogether/hedra-avatarand requiresstart_keyframe_idplus eitheraudio_idor anaudio_generationobject. - Auth is an
X-API-Keyheader onhttps://api.hedra.com/web-app/public. - The guide says the model catalog reports output limits separately from the limits on each input slot, and tells you to check them before uploading long audio.
| Step | Endpoint | Purpose |
|---|---|---|
| 1 | POST /assets, then POST /assets/{id}/upload | Create and fill the audio asset |
| 2 | POST /assets, then POST /assets/{id}/upload | Create and fill the portrait asset |
| 3 | POST /generations | Start the video from both assets |
| 4 | GET /generations/{id}/status | Poll progress until complete |
The Sume flow
Sume's avatar video guide takes exactly one of script or video_inputs, plus optional quality, aspect_ratio, scene and captions. Media fields must be fetchable public HTTPS URLs, so there is no separate asset-create step on Sume's side for those.
Before that you create the avatar itself once with POST /v1/avatar-1.0/generate (prompt, profile or a reference image), which is its own job. After that, each video is one POST.
- Create the avatar once, reuse the handle.
- POST the video with an
Idempotency-Key. - Poll
GET /v1/jobs/{id}/statusor receive a signedjob.completedwebhook; see Webhooks.
What you give up on Sume
Sume does not take your own finished audio on this route; it speaks the script. If your audio already exists, the Fabric route (POST /v1/veed/fabric-1.0) takes an audio_url and a measured duration_seconds with a still, as listed on the Models page.
Hedra's flow is more steps but fits recorded audio of any voice. Sume's is shorter but speaks 4 to 60 seconds per job.
Sources
Related posts
More in Developers
- HeyGen bulk status for 100 videos at once vs Sume job webhooks
HeyGen added bulk status reads of up to 100 ids for videos, lipsyncs, translations and assets. Sume reads jobs one by one or pushes signed terminal webhooks.
- HeyGen POST /v3/templates makes a template from video; Sume has none
HeyGen's October API entry turns a video into a private template. Sume has no template object for avatar video: you keep the request JSON and avatar handle.
- HeyGen lipsync batches of 100 vs Sume Fabric one clip per job
HeyGen added POST /v3/lipsyncs/batches. Sume's talking-still route, VEED Fabric 1.0, has no batch route: you submit one job per clip with idempotency keys.
- HeyGen Studio video scene voiceover: freeze, loop or fit_to_scene
HeyGen's Studio API lets a video scene carry voiceover with playback freeze, loop or fit_to_scene. Sume uses silence beats and a timeline instead.
Written by Sume