Make a portrait painting talk: H3 Max lip sync, 12 s for $1.20

A museum label, a family portrait or a book cover can speak: H3 Max lip sync takes a still and 5 to 14.8 seconds of audio. At 768p a 12 second clip is $1.20.

5 min readSume
All posts

The short answer

To make a portrait painting talk, send the painting as image_url and a Sume-hosted audio file to POST /v1/minimax/h3-max/lip-sync. Audio must run 5 to 14.8 seconds. At the default 768p the price is $0.10 per second rounded up, so an 11.4 second line reserves 12 seconds and costs $1.20.

What this route does

MiniMax H3 Max lip sync is a Sume surface for a still plus audio. It is not the prompt-driven minimax-h3-max video model: you do not write a scene, and the audio sets the length of the clip. The image must have an aspect ratio between 0.4 and 2.5, which covers most portrait paintings and framed photos.

Price for a 12 second line

Sume prices it as the fal per-second lip-sync list times the 1.25 house margin, on ceil(duration_seconds).

Billed per second, rounded up, read from docs/api/minimax-h3-max-lip-sync.md on 2026-10-05
ResolutionPer second12 s clip
480p$0.0625$0.75
768p (default)$0.10$1.20
1080p$0.20$2.40

Request

First, get the voice. Generate it with Sume text-to-speech, or import a recording through POST /v1/media-imports, so the audio_url sits on the Sume media host. Then submit. duration_seconds is required and only sets the reservation, so give the real length of the audio.

curl -X POST https://api.sume.com/v1/minimax/h3-max/lip-sync \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: portrait-001" \
  -d '{
    "image_url": "https://example.com/portrait.jpg",
    "audio_url": "https://media.sume.com/artifacts/artf_demo/line.wav",
    "duration_seconds": 11.4,
    "resolution": "768p",
    "mode": "async"
  }'

Limits to plan around

The route is strict about the audio window. Below 5 seconds or above 14.8 seconds the API returns invalid_request and does not clamp, because the provider rejects short audio and silently drops anything past 14.8 seconds, which would cut your sentence. A 20 second script needs to be split into two audio files first; Fabric (POST /v1/veed/fabric-1.0) stays the default talking-photo model and takes the same body.

Rights and labels

Only portraits you have the right to use should go in. A painting in the public domain, your own artwork, or a family portrait with permission are fine starting points. Label the result as AI-generated where the platform asks for it. For a more complete talking-photo recipe, see Make a photo talk with AI.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume