LTX-2.5 or a lip-sync API for a talking head: what Sume offers
Sume does not list LTX-2.5. For a talking head that must say your exact words, the route that ships is TTS plus H3 Max lip sync, or Avatar 1.0 talking video.

Sume does not list LTX-2.5, so there is no LTX-2.5 model id to call on Sume. A release tracker records LTX-2.5 as released on 2026-08-11 (Magic Hour tracker, read 2026-10-06), and that is the extent of what this page claims about it. For a talking head whose words must match a script exactly, the route Sume ships is a voice line from Sume TTS fed to the H3 Max lip-sync model, or the Avatar 1.0 talking-video route that does both steps.
The reason is not a quality ranking. It is a control question: a video model that generates its own sound decides what is said, while a lip-sync model is given the audio and moves the mouth to it.
What decides between the two approaches?
| Question | Video model with its own audio | Lip-sync model on Sume |
|---|---|---|
| Who writes the spoken words? | The prompt and the model | You, as audio_url |
| Can you re-record one line? | Re-roll the whole clip | Change the audio, keep the still |
| Is the face fixed? | Described in a prompt | Your image_url or avatar handle |
| Is it on Sume? | LTX-2.5: not listed | Yes, minimax/h3-max/lip-sync |
| Length window | Per model | 5 to 14.8 s of audio |
What would you actually do on Sume?
First, decide whether the face must be the same person across clips. If so, create an avatar once and use its handle, so each clip starts from an identical face. Second, produce the audio with Sume TTS and read its duration. Third, call the lip-sync route with the audio and the handle. If a model in the video catalog suits a scene without speech, such as a b-roll cutaway, use generate_video for that and cut it in; the catalog lists ids such as seedance-2.5, kling-3 and wan-3.0.
Sume's docs are explicit that video models do not lip-sync to TTS, so narration laid over a generated clip will not match mouth movements. That is the fact that settles most talking-head questions, whichever model name is trending.
What if you need LTX specifically?
Then Sume is not the place to get it today, and nothing here should be read as a promise otherwise. The stored post on what Sume lists instead of an LTX id keeps the current list. If your need is a face that talks, the lip-sync route covers it at $0.0625, $0.10 or $0.20 a second at 480p, 768p and 1080p; the rate comparison puts those next to Fabric. Check the catalog again before building, because it changes.
How do you test a talking head cheaply?
Write a single 8-second line, produce its audio, and run it once at 480p, which costs 8 x $0.0625 = $0.50. If the mouth shapes and the face look right, step up only for the placements that need it. If they do not, you have spent half a dollar to learn that, not a campaign budget.
Save the still, the audio and the handle with the output. When a trending model later appears in Sume's catalog, you can rerun the same inputs through it and compare directly, instead of arguing from release notes.
Sources
Related posts
More in Comparisons
- MAI-Transcribe-2 at $0.54 an hour vs Sume STT at 1 cent a minute
MAI-Transcribe-2-Streaming is $0.54 an hour as an intro price through 2026. Sume STT is $0.01 a minute, $0.60 an hour. The bill at 10, 100 and 1,000 hours.
- MAI-Voice-2.1 or Sume TTS? Three questions that decide it
Live speech, extra outputs or lowest price per character? A short guide to MAI-Voice-2.1 and Flash against the Sume TTS Router, with 10M-character math.
- MiniMax H3 open weights vs a hosted lip-sync API: what you take on
MiniMax released H3 open weights on 2026-08-03. Self-hosting is not the same as a hosted still-plus-audio lip-sync route. What each choice makes you own.
- Nova Canvas 4,194,304 pixel cap and 16 px rule vs Sume image_size
Nova Canvas output sides must divide by 16 and total under 4,194,304 pixels. GPT Image 2.5 on Sume has its own custom-size rules. A side-by-side check.
Written by Sume