LTX-2.5 or a lip-sync API for a talking head: what Sume offers

Sume does not list LTX-2.5. For a talking head that must say your exact words, the route that ships is TTS plus H3 Max lip sync, or Avatar 1.0 talking video.

5 min readSume
All posts

Sume does not list LTX-2.5, so there is no LTX-2.5 model id to call on Sume. A release tracker records LTX-2.5 as released on 2026-08-11 (Magic Hour tracker, read 2026-10-06), and that is the extent of what this page claims about it. For a talking head whose words must match a script exactly, the route Sume ships is a voice line from Sume TTS fed to the H3 Max lip-sync model, or the Avatar 1.0 talking-video route that does both steps.

The reason is not a quality ranking. It is a control question: a video model that generates its own sound decides what is said, while a lip-sync model is given the audio and moves the mouth to it.

What decides between the two approaches?

Approach comparison. Sume routes from docs.sume.com Models, read 2026-10-06; LTX-2.5 date from the Magic Hour tracker, read 2026-10-06.
QuestionVideo model with its own audioLip-sync model on Sume
Who writes the spoken words?The prompt and the modelYou, as audio_url
Can you re-record one line?Re-roll the whole clipChange the audio, keep the still
Is the face fixed?Described in a promptYour image_url or avatar handle
Is it on Sume?LTX-2.5: not listedYes, minimax/h3-max/lip-sync
Length windowPer model5 to 14.8 s of audio

What would you actually do on Sume?

First, decide whether the face must be the same person across clips. If so, create an avatar once and use its handle, so each clip starts from an identical face. Second, produce the audio with Sume TTS and read its duration. Third, call the lip-sync route with the audio and the handle. If a model in the video catalog suits a scene without speech, such as a b-roll cutaway, use generate_video for that and cut it in; the catalog lists ids such as seedance-2.5, kling-3 and wan-3.0.

Sume's docs are explicit that video models do not lip-sync to TTS, so narration laid over a generated clip will not match mouth movements. That is the fact that settles most talking-head questions, whichever model name is trending.

What if you need LTX specifically?

Then Sume is not the place to get it today, and nothing here should be read as a promise otherwise. The stored post on what Sume lists instead of an LTX id keeps the current list. If your need is a face that talks, the lip-sync route covers it at $0.0625, $0.10 or $0.20 a second at 480p, 768p and 1080p; the rate comparison puts those next to Fabric. Check the catalog again before building, because it changes.

How do you test a talking head cheaply?

Write a single 8-second line, produce its audio, and run it once at 480p, which costs 8 x $0.0625 = $0.50. If the mouth shapes and the face look right, step up only for the placements that need it. If they do not, you have spent half a dollar to learn that, not a campaign budget.

Save the still, the audio and the handle with the output. When a trending model later appears in Sume's catalog, you can rerun the same inputs through it and compare directly, instead of arguing from release notes.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume