InfiniteTalk's 40 s single clip vs a 60 s Sume avatar video
InfiniteTalk makes about 40 s in single-clip mode and unlimited length when streaming. Compare that with a Sume avatar video of 4 to 60 s, priced per second.
InfiniteTalk is an Apache 2.0 audio-driven video model that talks about "unlimited-length" generation. Its README lists 480P and 720P output, and says single-clip mode tops out near 40 seconds (about 1000 frames), while streaming mode has no fixed cap. It is built on Wan2.1-I2V-14B and supports more than one person in a shot.
If your clip is a single presenter and under a minute, you may not need that headroom. I read the InfiniteTalk README on 2026-10-04 and compare it with the Sume avatar video limits in the docs.
What the two systems limit
| Topic | InfiniteTalk | Sume avatar video |
|---|---|---|
| Length | About 40 s single clip; streaming unlimited | 4 to 60 s per request |
| Resolution | 480P or 720P | 720p |
| People per shot | Multi-person supported | One avatar per final video |
| Where it runs | Your hardware | Sume job, billed per second |
| Input | Image or video plus audio | Script, or scene list, plus an avatar |
When 60 seconds is enough
Most short-form placements sit inside 60 seconds, which is the ceiling for POST /v1/avatar-1.0/talking-video. You send a script (or a video_inputs scene list), and Sume estimates the spoken duration from the text. If the estimate falls outside 4 to 60 seconds the request is rejected, so you learn the problem before spending. Quality is standard, plus (the default) or max, and aspect ratio is 1:1, 3:4, 9:16 (the default), 4:3 or 16:9.
What it costs per second
Rates with no product image are $0.184 per second at standard, $0.245 at plus and $0.55 at max. A full 60 second clip at plus is therefore 60 x $0.245 = $14.70. If the script is long, split it into two requests and keep the same avatar handle, so the face stays identical between parts.
- Standard: $0.184 per second, 60 s = $11.04.
- Plus: $0.245 per second, 60 s = $14.70.
- Max: $0.55 per second, 60 s = $33.00.
When InfiniteTalk is the better fit
Choose a self-run model when you need several people talking in one frame, when a single take must run past a minute, or when you already hold the GPU capacity. Choose Sume when the shot is one presenter, the script is under 60 seconds, and you would rather approve a still preview first: previews are independent of quality tier, so you can approve once and then pick the tier at generation time.
Sources
Related posts
More in Sume Avatar 1.0
- Italy deepfake offense: 1-5 years, and avatar consent records
A secondary source says Italy's Law 132/2025 took effect Oct 10, 2025 and Art. 612-quater carries 1-5 years for deepfakes. Keep consent records for avatars.
- LatentSync 1.6 needs 18 GB VRAM: a hosted lip sync from a still
LatentSync 1.6 is a 512 px video-plus-audio lip-sync model that wants 18 GB of GPU memory. No GPU, only a photo? Here is the Sume still-plus-audio route.
- LivePortrait needs a driving video, not audio: Kling motion control
LivePortrait animates a portrait from a driving video, not from speech. Sume's Kling 3.0 motion control route does that as a hosted API, per second.
- Nonprofit donor thank-you: three tiered avatar videos, total cost
Three 20-second thank-you videos, one per giving tier, cost $11.04 in total at standard. See how to script them and why you review a preview first.
Written by Sume