LTX-2.5 audio-to-video: three ways to drive a clip with audio on Sume
LTX-2.5 is not on Sume. To make video that follows a voice or track, use audio references on Seedance 2.5, Wan 3.0 or H3, or H3 Max lip sync. Limits and prices.

LTX-2.5 (Lightricks, 2026-08-11) is listed on Magic Hour's tracker, read 2026-10-06, as going up to 4K with an audio-to-video mode. It is not in Sume's video catalog. If what you want is video that follows a voice or a music track, Sume has three routes: audio references on seedance-2.5, wan-3.0 and minimax-h3; a lip-sync job on MiniMax H3 Max for a still image plus a voice; and text to speech first, then either of the above.
The vendor claim is on the tracker, and the Sume limits are in the Video Router docs and the lip-sync spec in the video docs.
Three routes compared
Pick by what you start with. A photo and a voice is lip sync. A prompt plus a track is an audio reference. A script and no voice is TTS first.
| Route | Starts from | Length | Price per second |
|---|---|---|---|
Audio reference on seedance-2.5 | prompt + reference_audio_urls | 4-30 s | from $0.2687 (480p) |
Audio reference on wan-3.0 | prompt + reference_audio_urls | 2-30 s | $0.0625 to $0.25 |
Audio reference on minimax-h3 | image or video + reference_audio_urls | 5-15 s | $0.0625 to $0.075 |
minimax/h3-max/lip-sync | still image + audio | 5-14.8 s of audio | $0.0625 to $0.20 |
Audio reference limits
On Wan 3.0 you can send up to 5 audio clips totaling at most 15 seconds. On MiniMax H3 you can send up to 3, each 2 to 15 seconds and combined at most 15 seconds, and the audio cannot be the only reference. Seedance 2.5 accepts audio references too; check supported_input_references for the current list. In all three, the audio steers the clip, and the model decides how literally to follow it.
If you need lips to match a specific line, use the lip-sync route instead. The audio there sets the output length and the mouth movement, rather than serving as a hint.
The lip-sync route
POST /v1/minimax/h3-max/lip-sync takes a still image (or a ready avatar) and a Sume-hosted audio file of 5 to 14.8 seconds, up to 10 MB, and returns a talking clip at 480p, 768p or 1080p. A 5-second 768p clip reserves $0.50. Audio shorter than 5 seconds is refused, and the API never clamps the length for you, because silent clipping would drop part of the speech.
import os, requests
r = requests.post(
"https://api.sume.com/v1/minimax/h3-max/lip-sync",
headers={"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"},
json={
"image_url": "https://example.com/host.jpg",
"audio_url": "https://media.sume.com/example/line.mp3",
"duration_seconds": 6,
"resolution": "768p",
"mode": "async",
},
timeout=60,
)
print(r.status_code, r.json())How to choose
Use lip sync when the face is the subject and the words must match. Use an audio reference when the scene is the subject and the track sets its rhythm. Use TTS first when you have a script but no voice: generate the voice, check it, then feed it to either route. A 20-second voiceover does not fit the lip-sync window, so cut it into clips of 5 to 14.8 seconds, or use an audio reference on Wan or Seedance, which accept longer clips.
What Sume does not have is LTX's single audio-to-video mode at up to 4K. If that exact capability is the requirement, you would run LTX where it is offered. For the other cases, the audio-reference comparison and the catalog check are the next reads.
A worked example: a 12-second spoken ad
Say you have a 12-second voiceover and one product photo. The lip-sync route fits, since 12 seconds is inside the 5 to 14.8 second window. At 768p it reserves $1.20, and at 1080p $2.40. If the speaker is not on camera and you only want visuals that follow the line, an audio reference on Wan 3.0 at 720p costs $1.50 for the same length. The first spends more for exact mouth sync; the second spends less and loosens the match. Decide by whether a viewer will watch the lips.
Whichever you pick, test with a 5-second slice first. A slice costs a fraction of the full job, and mouth sync problems show up in the first two seconds. Fix the audio level and the framing there, then run the full clip.
What to check before you send
Host your audio on a URL the API can fetch, keep it under the 10 MB lip-sync limit, and confirm its length with a probe tool, because a 4.9-second file is refused. For audio references, confirm the combined length is at most 15 seconds on H3 and Wan. Then read the model's entry in /v1/videos/models so you are sending only values the catalog lists.
Sources
Related posts
More in Models
- MAI-Voice-2.1 vs Flash: what $7 per million saves on a 60-second short
Flash lists at $15 per million characters and MAI-Voice-2.1 at $22. On a 750-character short that is a fraction of a cent. Math for 1, 30 and 3,000 shorts.
- MiniMax H3 first request: a 5-second 768p clip for 38 cents on Sume
A 5-second 768p MiniMax H3 clip costs $0.38 on Sume, with stereo sound included. The request, the 15-second limit and the 2K and 4K upscale prices.
- MiniMax H3 reference limits: 9 images, 3 videos, 3 audios on Sume
MiniMax H3 reference-to-video on Sume takes up to 9 images, 3 videos and 3 audio clips, 12 in total. The duration rules, the extra-image fee and a request.
- Native 4K AI video API: LTX-2.5 claims 4K, what Sume offers instead
LTX-2.5 is listed up to 4K but is not in Sume's catalog. Sume's 4K options are Gemini Omni Flash 1.1 and MiniMax H3 4K upscales. Limits and per-second prices.
Written by Sume