Lip-sync a saved avatar with H3 Max: avatar_handle, not image_url
Use a ready Sume avatar as the face for MiniMax H3 Max Lip Sync: send avatar_handle plus Sume-hosted audio, 5-14.8 seconds, and see the price at 768p.
To lip-sync a saved avatar with MiniMax H3 Max, post to POST /v1/minimax/h3-max/lip-sync with avatar_handle instead of image_url, a Sume-hosted audio_url, and a duration_seconds between 5 and 14.8. A 10-second line at 768p reserves $1.00 on Sume. The body is the same as VEED Fabric 1.0: exactly one visual source, either image_url or an avatar id or handle, never both. Sume lists the route in its models docs.
Why use the handle
A handle means the same face in every clip without hosting a still. Use a handle whose avatar job has finished. The audio must live on Sume's media host and be at most 10 MB, so run speech through Sume TTS first or upload through Sume's media inputs.
resolutionis480p,768p(default) or1080p; there is no 2K.speed_tieris accepted and ignored.- The API rejects
modeland provider endpoint fields: the URL is the model.
Price at each resolution
Sume's rate is the fal list price times 1.25. fal's own page lists $0.05, $0.08 and $0.16 per second at 480p, 768p and 1080p; Sume's per-second rates are $0.0625, $0.10 and $0.20.
| Resolution | Sume USD per second | 10-second clip |
|---|---|---|
| 480p | 0.0625 | $0.63 |
| 768p | 0.10 | $1.00 |
| 1080p | 0.20 | $2.00 |
Request
Replace the audio URL with your own Sume-hosted file.
curl -X POST https://api.sume.com/v1/minimax/h3-max/lip-sync \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: h3-handle-001" \
-d '{
"avatar_handle": "support_host",
"audio_url": "https://media.sume.com/audio/line.mp3",
"duration_seconds": 10,
"resolution": "768p"
}'
curl https://api.sume.com/v1/jobs/JOB_ID/status -H "Authorization: Bearer $SUME_API_KEY"Sources
Related posts
More in Media tools
- Put a logo or lower third over a talking avatar video with compose
Timeline compose overlay places a Sume-hosted still over an avatar clip, at the top, center or bottom. $0.02 flat per job. The layout keys and a request.
- Use a MAI-Voice-2.1 clip as avatar audio? Sume needs its own file
A talking still on Sume takes Sume-hosted audio under 10 MB. A MAI-Voice-2.1 clip sits elsewhere, so make the line with Sume TTS instead. Sizes and limits.
- Match voice emotion and music mood: one mood word, two Sume fields
Set generation_config.emotion on TTS and the emotion axis in the Music prompt from the same mood word, so a short's voice and bed do not argue.
- Microsoft's Content Provenance Detection: what you can check on a file
Foundry has a detection website and API for provenance. What the page says it checks, its limits, and how to use it on an AI clip or image from any generator.
Written by Sume