Voice greeting on a still photo: TTS first, then H3 Max lip sync
To make a still photo speak, make 5 to 14.8 seconds of TTS audio, then call MiniMax H3 Max lip sync. TTS costs $0.0475 per 1,000 characters.

A still photo can speak a greeting if you do it in two calls: generate the voice with TTS, then send the photo and the audio to POST /v1/minimax/h3-max/lip-sync. The audio must be 5 to 14.8 seconds, and TTS costs $0.0475 per 1,000 characters, so a 400-character greeting is $0.019 of voice before the lip-sync charge.
Step 1: the voice
Use POST /v1/tts-1.0/generate with a transcript and a voice selector (avatar_handle or voice.id). Ask for wav and measure the duration of the file you get back, because the lip-sync call needs duration_seconds.
| Characters | Rate | Cost |
|---|---|---|
| 120 | $0.0475 per 1,000 | $0.0057 |
| 400 | $0.0475 per 1,000 | $0.019 |
| 1,200 | $0.0475 per 1,000 | $0.057 |
Step 2: the lip-sync call
The body takes exactly one visual source (image_url or an avatar handle) plus a Sume-hosted audio_url and duration_seconds. The default resolution is 768p; 480p and 1080p are also allowed. Sume reserves the list price times 1.25 at admit, so read the live figure in GET /v1/catalog rather than assuming one.
curl -X POST https://api.sume.com/v1/minimax/h3-max/lip-sync \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: greeting-photo-001" \
-d '{
"image_url": "https://example.com/portrait.png",
"audio_url": "https://media.sume.com/artifacts/artf_demo/greeting.wav",
"duration_seconds": 9.4,
"resolution": "768p"
}'Limits and gotchas
- Audio under 5 seconds or over 14.8 seconds is refused, not clamped. Split a longer script into several sentences and several calls.
audio_urlmust be on the Sume media host and no larger than 10 MB; the TTS output already is.- The still must have an aspect ratio between 0.4 and 2.5.
- Never put someone else's face on a script they did not approve.
Sources
Related posts
More in Media tools
- Voiceover plus music bed: TTS and Lyria cost, duck_db in Timeline
A 700-character voiceover ($0.03325) and one Lyria bed ($0.125) cost $0.15825 before the render. Timeline duck_db (0 to 20) lowers the bed under speech.
- Which video specs TikTok, LinkedIn, Pinterest and Shorts leave blank
Only TikTok lists a bitrate; only LinkedIn lists a frame rate; Pinterest lists no resolution; the Shorts page lists no size at all. What Sume can set for each.
- X video ad thumbnail: PNG or JPEG under 5 MB, pulled from your own MP4
X accepts a PNG or JPEG thumbnail up to 5 MB that matches the video. Sume video frames returns a still at a time you pick, at source size.
- X website-button video: headline sits 8-12 px in, move captions
X's spec puts the headline 8 to 12 pixels from the left and bottom of a website-button video. Keep burned-in Sume captions clear of that corner at 800x450.
Written by Sume