Make a portrait painting talk: H3 Max lip sync, 12 s for $1.20
A museum label, a family portrait or a book cover can speak: H3 Max lip sync takes a still and 5 to 14.8 seconds of audio. At 768p a 12 second clip is $1.20.

The short answer
To make a portrait painting talk, send the painting as image_url and a Sume-hosted audio file to POST /v1/minimax/h3-max/lip-sync. Audio must run 5 to 14.8 seconds. At the default 768p the price is $0.10 per second rounded up, so an 11.4 second line reserves 12 seconds and costs $1.20.
What this route does
MiniMax H3 Max lip sync is a Sume surface for a still plus audio. It is not the prompt-driven minimax-h3-max video model: you do not write a scene, and the audio sets the length of the clip. The image must have an aspect ratio between 0.4 and 2.5, which covers most portrait paintings and framed photos.
Price for a 12 second line
Sume prices it as the fal per-second lip-sync list times the 1.25 house margin, on ceil(duration_seconds).
| Resolution | Per second | 12 s clip |
|---|---|---|
| 480p | $0.0625 | $0.75 |
| 768p (default) | $0.10 | $1.20 |
| 1080p | $0.20 | $2.40 |
Request
First, get the voice. Generate it with Sume text-to-speech, or import a recording through POST /v1/media-imports, so the audio_url sits on the Sume media host. Then submit. duration_seconds is required and only sets the reservation, so give the real length of the audio.
curl -X POST https://api.sume.com/v1/minimax/h3-max/lip-sync \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: portrait-001" \
-d '{
"image_url": "https://example.com/portrait.jpg",
"audio_url": "https://media.sume.com/artifacts/artf_demo/line.wav",
"duration_seconds": 11.4,
"resolution": "768p",
"mode": "async"
}'Limits to plan around
The route is strict about the audio window. Below 5 seconds or above 14.8 seconds the API returns invalid_request and does not clamp, because the provider rejects short audio and silently drops anything past 14.8 seconds, which would cut your sentence. A 20 second script needs to be split into two audio files first; Fabric (POST /v1/veed/fabric-1.0) stays the default talking-photo model and takes the same body.
Rights and labels
Only portraits you have the right to use should go in. A painting in the public domain, your own artwork, or a family portrait with permission are fine starting points. Label the result as AI-generated where the platform asks for it. For a more complete talking-photo recipe, see Make a photo talk with AI.
Sources
Related posts
More in Use cases
- Make an AI voice read an email or order code right: spell, verify
Amazon and Cartesia both claim better codes, emails and phone numbers. The reliable fix is text prep plus a read-back check. Try both on Sume's TTS router.
- Marketplace listing video from one product photo: first_frame
Turn a single product photo into a short listing video by pinning it as the first frame on seedance-2, then check the catalog fields that gate the request.
- Mascot dance ad: Kling motion control, orientation and sound
Drive a still mascot with a motion reference video through Kling 3.0 Motion Control on Sume. Settings that matter: duration, orientation and original sound.
- Mattress Black Friday video ad: a dimmed bedroom clip and a price card
Make a mattress Black Friday ad with AI: one 8-second bedroom clip, a dim pass, and a price-card overlay on top, about $1.12 in Sume jobs.
Written by Sume