Lip sync a Korean voice recording on Sume: which route, what it costs
Avatar Video is English only in code. For a Korean recording, Sume offers audio-driven routes: H3 Max lip-sync and VEED Fabric. Windows, rates and a test plan.
For a Korean recording, skip the script-driven Avatar Video route and use an audio-driven one: MiniMax H3 Max lip-sync (5 to 14.8 second audio, $0.10 per second at 768p) or VEED Fabric 1.0 ($0.1875 per second at 720p). Both take a still image or ready avatar plus Sume-hosted audio and make the mouth follow it. The talking-video route's clip prompt is English only, so it is the wrong tool.
| Route | Audio window | Rates per audio second | Note |
|---|---|---|---|
| minimax/h3-max/lip-sync | 5 to 14.8 s | 480p $0.0625, 768p $0.10, 1080p $0.20 | Output length follows audio |
| veed/fabric-1.0 | Estimated from 1 s; maximum priced at 300 s | 480p $0.10, 720p $0.1875 | Primary Fabric route |
| sume/avatar-1.0/fabric | Capped at 15 s | 720p $0.3024 per output second | Experimental, test-only, id will change |
Preparing the recording
Both routes need audio Sume can fetch, which in practice means a Sume-hosted URL. If your Korean voice is inside a video, extract it first with audio detach, which writes a wav by default. The timeline audio page notes that wav is the format to keep when a file drives lip sync, because mp3 adds priming padding at every edge. Audio detach is a flat $0.01 per job, per the OpenAPI text for that route.
Cut the audio to the window. For H3 Max that is 5 to 14.8 seconds, so a 40 second recording becomes three or four segments.
Cost of a 12 second line
For a single 12 second Korean line:
- H3 Max at 768p: 12 x $0.10 = $1.20.
- H3 Max at 1080p: 12 x $0.20 = $2.40.
- Fabric at 720p: 12 x $0.1875 = $2.25.
- Fabric at 480p: 12 x $0.10 = $1.20.
What Sume does not claim
The route descriptions say the mouth follows the audio. They do not publish a language list or a lip-sync accuracy figure, so the Korean result is something you verify. Make one 5 second test at H3 Max 480p, which is 5 x $0.0625 = $0.3125, and watch the mouth against the audio with the sound on. If it passes, scale up. Add Hangul subtitles with the korean-ad or another Hangul style through Video captions.
Which route to pick
Use H3 Max when the audio fits 5 to 14.8 seconds and you want the lower rate. Use Fabric when the recording is longer than that window, since its priced range runs up to 300 seconds, but note it costs $0.1875 per second at 720p. Do not use the experimental sume/avatar-1.0/fabric route in production: its description says it is test-only and that the model id will change.
Whichever you choose, send exactly one visual source: either an image_url of a posed still or an avatar_handle, never both.
Sources
Related posts
More in Sume Avatar 1.0
- Live AI avatar vs recorded clip: who reviews the words first?
A live avatar speaks in real time; a Sume Avatar 1.0 clip is scripted, previewable and fixed. Why that matters while Tavus says disclosure features are coming.
- Longest Avatar 1.0 video is 60 seconds: $11.04 on standard
A 60-second Avatar 1.0 talking video costs $11.04 on standard, $14.70 on plus and $33.00 on max. The cap, the product-image rate and a split plan.
- Make an AI avatar from a prompt, traits or a photo: $0.95 a call
Sume Avatar 1.0 creates a reusable avatar from a text prompt, structured traits or a reference image. All three cost the same flat $0.95. See the bodies.
- Make an AI avatar from a selfie: the photo URL rules on Sume
Sume builds an avatar from a prompt, a profile or a photo. The photo must be a public HTTPS image URL; localhost, private and non-image URLs are rejected.
Written by Sume