Audio to avatar AI: make an avatar speak your recording
Audio to avatar AI lip-syncs a face to a voice track you supply. How it works, which APIs take a recording directly, and the route on Sume today.
Audio to avatar AI takes a voice recording and makes a face, either a photo or a ready-made avatar, speak it, with the mouth moving in time to the audio. It is lip sync driven by your audio instead of a typed script, and the video runs as long as the recording. HeyGen's API, for example, takes an audio_url in place of a script.
On Sume, the lip-sync models accept only audio already on Sume's media host, so a recording made elsewhere can't go in directly today. The route is to clone your voice once, have text to speech read the script in it, and lip-sync that audio. Sume facts come from the Models overview and the Sume API reference, read on 2026-09-28; the Voices steps are read from the Sume app's current code.
How does audio to avatar work?
- You supply a face: a single photo of one person, or an avatar the tool already has.
- You supply the audio: your recording, a voice-over from another tool, or a dubbed track.
- A lip sync model redraws the mouth and face to match the audio, and the clip's length follows the audio.
- Script-to-avatar is the other mode: you send text and the tool generates the voice as well as the video.
Which APIs take a recording directly?
HeyGen's Audio to Video takes a public HTTPS audio_url or an uploaded MP3 or WAV of up to 32 MB, and renders up to 30 minutes per request onto an avatar or a single-person image. Creatify's Aurora takes one photo and an mp3 or wav URL of up to 5 minutes, and its AI Avatar endpoint takes your audio URL instead of text. Lip sync AI API options compares their limits and prices.
Can I use my own audio file on Sume?
Not directly today. Each Sume route that makes a face talk either refuses audio from outside Sume or takes no audio at all:
| Sume route | What it does with audio |
|---|---|
VEED Fabric 1.0, POST /v1/veed/fabric-1.0 | Lip-syncs a still or a ready avatar to an audio_url on the Sume media host, typically a TTS segment, at most 10 MB. Other hosts are rejected. |
MiniMax H3 Max Lip Sync, POST /v1/minimax/h3-max/lip-sync | The same still-plus-audio body as Fabric, for 5–14.8 seconds of audio |
Avatar 1.0 talking video, POST /v1/avatar-1.0/talking-video | Takes a script or per-scene text, not a voice track |
| Asset upload | Implemented but hidden from the public OpenAPI, so the public API has no route that puts a local recording on the media host |
How do I make an avatar speak in my voice on Sume?
Use your recording to clone the voice rather than as the soundtrack. Clone it once in the Sume app under Assets → Voices, from a clip you upload or record there. Then POST /v1/tts-1.0/generate speaks your script in that voice, with its voi_ id as voice.id, and POST /v1/veed/fabric-1.0 lip-syncs a photo or a ready avatar to the TTS audio. The avatar says your script in your cloned voice; it does not replay the original recording.
Text to speech costs $0.0475 per 1,000 characters, and Fabric costs $0.1875 per audio second (720p), rounded up to the whole second, each plus a 5.5% agent fee by default. How to clone yourself with AI has both requests and their limits, and AI voiceover in your own voice walks through the cloning step.
Sources
Related posts
More in Sume Avatar 1.0
- How to create an avatar from a video: two ways
To create an avatar from a video, train a digital twin on the footage, or take one clear frame and make a photo avatar from it. What each route keeps.
- AI avatar video editing: what you can change after rendering
AI avatar video editing: trims, captions, music, B-roll, and joins work on the finished MP4; new words, a new voice, or a new look need a re-render.
- HeyGen Avatar 4 vs 5: Avatar IV and Avatar V compared
HeyGen Avatar IV is the default engine for every avatar type; Avatar V is opt-in for eligible Digital Twins only. Parameters, eligibility and cost.
- HeyGen video translation API: modes, languages, cost
HeyGen's video translation API is POST /v3/video-translations: a video URL plus target languages, in Speed or Precision mode, billed per minute.
Written by Sume