Talking avatar from an approved TTS file: image-to-video audio_url

To animate an avatar from a voiceover you already approved, send its Sume-hosted audio_url and duration_seconds to Avatar 1.0 image-to-video. Limits inside.

5 min readSume
All posts

To make an avatar speak a voiceover you have already approved, send the TTS file to POST /v1/avatar-1.0/image-to-video as audio_url, with an avatar_handle (or avatar_id, or an image_url) and a duration_seconds. The audio has to be a Sume-hosted file of at most 10 MB, and the video length follows the audio.

Fields are from the Avatar image-to-video request in the Sume API reference (read 2026-10-06).

What does the request need?

Exactly one visual source: image_url or avatar_id / avatar_handle, plus audio_url on the Sume media host and duration_seconds from 1 to 300, which Sume uses to reserve credits at admit. resolution is 480p or 720p, default 720p.

curl -X POST https://api.sume.com/v1/avatar-1.0/image-to-video \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: approved-vo-001" \
  -d '{
    "avatar_handle": "product_host",
    "audio_url": "https://media.sume.com/artifacts/artf_demo/voiceover.wav",
    "duration_seconds": 12,
    "resolution": "720p"
  }'

Why start from a file?

With talking-video you give a script and Sume voices it. With image-to-video you give the finished audio, so a take you approved by ear is the take that gets lip-synced. Poll the job as usual with GET /v1/jobs/{id}/status, as in Jobs and results.

What goes wrong?

A non-Sume audio host is rejected, and a file over 10 MB fails. Generate the TTS as WAV, keep the line short enough for the clip you want, and read the real length from the TTS result rather than guessing duration_seconds.

Sources

More in Sume Avatar 1.0

All Sume Avatar 1.0 posts

Written by Sume