LTX-2.5 audio-to-video: three ways to drive a clip with audio on Sume

LTX-2.5 is not on Sume. To make video that follows a voice or track, use audio references on Seedance 2.5, Wan 3.0 or H3, or H3 Max lip sync. Limits and prices.

6 min readSume
All posts

LTX-2.5 (Lightricks, 2026-08-11) is listed on Magic Hour's tracker, read 2026-10-06, as going up to 4K with an audio-to-video mode. It is not in Sume's video catalog. If what you want is video that follows a voice or a music track, Sume has three routes: audio references on seedance-2.5, wan-3.0 and minimax-h3; a lip-sync job on MiniMax H3 Max for a still image plus a voice; and text to speech first, then either of the above.

The vendor claim is on the tracker, and the Sume limits are in the Video Router docs and the lip-sync spec in the video docs.

Three routes compared

Pick by what you start with. A photo and a voice is lip sync. A prompt plus a track is an audio reference. A script and no voice is TTS first.

Limits from the Sume Video Router docs and the H3 Max lip-sync spec; prices are billed (list x 1.25), read 2026-10-06.
RouteStarts fromLengthPrice per second
Audio reference on seedance-2.5prompt + reference_audio_urls4-30 sfrom $0.2687 (480p)
Audio reference on wan-3.0prompt + reference_audio_urls2-30 s$0.0625 to $0.25
Audio reference on minimax-h3image or video + reference_audio_urls5-15 s$0.0625 to $0.075
minimax/h3-max/lip-syncstill image + audio5-14.8 s of audio$0.0625 to $0.20

Audio reference limits

On Wan 3.0 you can send up to 5 audio clips totaling at most 15 seconds. On MiniMax H3 you can send up to 3, each 2 to 15 seconds and combined at most 15 seconds, and the audio cannot be the only reference. Seedance 2.5 accepts audio references too; check supported_input_references for the current list. In all three, the audio steers the clip, and the model decides how literally to follow it.

If you need lips to match a specific line, use the lip-sync route instead. The audio there sets the output length and the mouth movement, rather than serving as a hint.

The lip-sync route

POST /v1/minimax/h3-max/lip-sync takes a still image (or a ready avatar) and a Sume-hosted audio file of 5 to 14.8 seconds, up to 10 MB, and returns a talking clip at 480p, 768p or 1080p. A 5-second 768p clip reserves $0.50. Audio shorter than 5 seconds is refused, and the API never clamps the length for you, because silent clipping would drop part of the speech.

import os, requests

r = requests.post(
    "https://api.sume.com/v1/minimax/h3-max/lip-sync",
    headers={"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"},
    json={
        "image_url": "https://example.com/host.jpg",
        "audio_url": "https://media.sume.com/example/line.mp3",
        "duration_seconds": 6,
        "resolution": "768p",
        "mode": "async",
    },
    timeout=60,
)
print(r.status_code, r.json())

How to choose

Use lip sync when the face is the subject and the words must match. Use an audio reference when the scene is the subject and the track sets its rhythm. Use TTS first when you have a script but no voice: generate the voice, check it, then feed it to either route. A 20-second voiceover does not fit the lip-sync window, so cut it into clips of 5 to 14.8 seconds, or use an audio reference on Wan or Seedance, which accept longer clips.

What Sume does not have is LTX's single audio-to-video mode at up to 4K. If that exact capability is the requirement, you would run LTX where it is offered. For the other cases, the audio-reference comparison and the catalog check are the next reads.

A worked example: a 12-second spoken ad

Say you have a 12-second voiceover and one product photo. The lip-sync route fits, since 12 seconds is inside the 5 to 14.8 second window. At 768p it reserves $1.20, and at 1080p $2.40. If the speaker is not on camera and you only want visuals that follow the line, an audio reference on Wan 3.0 at 720p costs $1.50 for the same length. The first spends more for exact mouth sync; the second spends less and loosens the match. Decide by whether a viewer will watch the lips.

Whichever you pick, test with a 5-second slice first. A slice costs a fraction of the full job, and mouth sync problems show up in the first two seconds. Fix the audio level and the framing there, then run the full clip.

What to check before you send

Host your audio on a URL the API can fetch, keep it under the 10 MB lip-sync limit, and confirm its length with a probe tool, because a 4.9-second file is refused. For audio references, confirm the combined length is at most 15 seconds on H3 and Wan. Then read the model's entry in /v1/videos/models so you are sending only values the catalog lists.

Sources

Related posts

More in Models

All Models posts

Written by Sume