Argil 500-character moment limit vs Sume's 4-60 second script window

Argil caps each moment transcript at 500 characters; Sume sizes a script by estimated seconds instead. How to split a long avatar script on either one.

5 min readSume
All posts

Argil limits every moment in a video to 500 characters of transcript, so a long script has to be cut into many moments. Sume does not count characters for a whole script: it accepts a script or multi-scene plan whose estimated duration lands between 4 and 60 seconds, and rejects anything outside that window.

Both limits come from vendor pages read on 2026-10-02: Argil's Create a new Video and Create a video pages, and Sume's Generate avatar video. This post is about how to plan a script against each limit, not about which avatar looks better.

What exactly does Argil limit per moment?

On the create-video endpoint, moments is a required array and each moment needs an avatarId plus either a transcript or an audioUrl, never both. The transcript of a single moment is capped at 500 characters. Argil's guide adds that a video can hold 60 paragraphs of up to 500 characters each, which it says can reach 10-15 minutes depending on how fast the avatar speaks.

The practical effect is that you chunk by characters. A 2,400-character narration becomes at least five moments, and you decide where the breaks fall.

How does Sume size a script instead?

Sume takes either a script or ordered video_inputs on POST /v1/avatar-1.0/talking-video, and exactly one of them. It estimates the spoken duration and accepts the request only if that estimate is 4-60 seconds inclusive. For a longer piece, the docs say to shorten the script or split it into multiple jobs.

Multi-scene plans use the same window for the total planned duration. Each spoken scene sets voice.type: "text" with a script or input_text and a duration, and a voice.type: "silence" scene is a non-speaking beat that needs a duration and no text.

Side by side

The two limits measure different things, so the comparison below is about units, not about which is larger.

Script limits as documented, read 2026-10-02
QuestionArgilSume
Unit of the limitCharacters per moment (500)Estimated seconds for the whole video (4-60)
Script or audiotranscript or audioUrl per momentscript or video_inputs, exactly one
Non-speaking beatNot documented on the pages readvoice.type: silence with a duration
Longer than the limitAdd moments; Argil guide cites 60 paragraphsSplit into several jobs
Avatar per videoavatarId on each momentOne resolved avatar per final video

How should I split a script for each?

For Argil, split on sentence boundaries and count characters before you send, so no moment is rejected. Keep each moment a complete thought, because the moment is also the unit you can attach gestures, zoom or b-roll to.

For Sume, estimate seconds rather than characters. A short product pitch fits one job. A longer explainer becomes several jobs that you submit with distinct Idempotency-Key headers and join afterwards. Sume's Jobs and results page says a queued job is a normal accepted state, so submit the pieces together and poll or use a webhook rather than waiting on one at a time.

What does Sume not do here?

Sume does not give you a single avatar video longer than 60 seconds, and the docs state that one final video uses one resolved avatar. If you need a ten-minute talking presenter in one render, Argil's moment model is built for that length and Sume's is not. The Sume route fits clips: hooks, product pitches, short explainers, and per-recipient variants. See AI avatar video longer than 60 seconds for the join-it-yourself approach.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume