Hedra legacy avatar API needs 4 calls; Sume avatar video needs 1 POST

Hedra's legacy avatar guide creates and uploads two assets, then generates and polls. Sume's talking-video route is one POST with a handle and a script.

5 min readSume
All posts

Hedra's legacy avatar guide has you upload audio, upload a portrait, start a generation with both asset ids, and poll until it completes. Sume's route is a single POST /v1/avatar-1.0/talking-video with an avatar_handle and a script, then polling the returned job.

The difference is where the inputs live. Hedra treats audio and image as files you place on its side first. Sume's request carries a reference to an avatar you created earlier and the text to speak.

The Hedra flow as documented

From the Hedra legacy avatar guide, read 2026-10-02:

  • The guide lists model_slug as together/hedra-avatar and requires start_keyframe_id plus either audio_id or an audio_generation object.
  • Auth is an X-API-Key header on https://api.hedra.com/web-app/public.
  • The guide says the model catalog reports output limits separately from the limits on each input slot, and tells you to check them before uploading long audio.
Hedra legacy avatar endpoints (read 2026-10-02)
StepEndpointPurpose
1POST /assets, then POST /assets/{id}/uploadCreate and fill the audio asset
2POST /assets, then POST /assets/{id}/uploadCreate and fill the portrait asset
3POST /generationsStart the video from both assets
4GET /generations/{id}/statusPoll progress until complete

The Sume flow

Sume's avatar video guide takes exactly one of script or video_inputs, plus optional quality, aspect_ratio, scene and captions. Media fields must be fetchable public HTTPS URLs, so there is no separate asset-create step on Sume's side for those.

Before that you create the avatar itself once with POST /v1/avatar-1.0/generate (prompt, profile or a reference image), which is its own job. After that, each video is one POST.

  • Create the avatar once, reuse the handle.
  • POST the video with an Idempotency-Key.
  • Poll GET /v1/jobs/{id}/status or receive a signed job.completed webhook; see Webhooks.

What you give up on Sume

Sume does not take your own finished audio on this route; it speaks the script. If your audio already exists, the Fabric route (POST /v1/veed/fabric-1.0) takes an audio_url and a measured duration_seconds with a still, as listed on the Models page.

Hedra's flow is more steps but fits recorded audio of any voice. Sume's is shorter but speaks 4 to 60 seconds per job.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume