Argil 500-character moment limit vs Sume's 4-60 second script window
Argil caps each moment transcript at 500 characters; Sume sizes a script by estimated seconds instead. How to split a long avatar script on either one.

Argil limits every moment in a video to 500 characters of transcript, so a long script has to be cut into many moments. Sume does not count characters for a whole script: it accepts a script or multi-scene plan whose estimated duration lands between 4 and 60 seconds, and rejects anything outside that window.
Both limits come from vendor pages read on 2026-10-02: Argil's Create a new Video and Create a video pages, and Sume's Generate avatar video. This post is about how to plan a script against each limit, not about which avatar looks better.
What exactly does Argil limit per moment?
On the create-video endpoint, moments is a required array and each moment needs an avatarId plus either a transcript or an audioUrl, never both. The transcript of a single moment is capped at 500 characters. Argil's guide adds that a video can hold 60 paragraphs of up to 500 characters each, which it says can reach 10-15 minutes depending on how fast the avatar speaks.
The practical effect is that you chunk by characters. A 2,400-character narration becomes at least five moments, and you decide where the breaks fall.
How does Sume size a script instead?
Sume takes either a script or ordered video_inputs on POST /v1/avatar-1.0/talking-video, and exactly one of them. It estimates the spoken duration and accepts the request only if that estimate is 4-60 seconds inclusive. For a longer piece, the docs say to shorten the script or split it into multiple jobs.
Multi-scene plans use the same window for the total planned duration. Each spoken scene sets voice.type: "text" with a script or input_text and a duration, and a voice.type: "silence" scene is a non-speaking beat that needs a duration and no text.
Side by side
The two limits measure different things, so the comparison below is about units, not about which is larger.
| Question | Argil | Sume |
|---|---|---|
| Unit of the limit | Characters per moment (500) | Estimated seconds for the whole video (4-60) |
| Script or audio | transcript or audioUrl per moment | script or video_inputs, exactly one |
| Non-speaking beat | Not documented on the pages read | voice.type: silence with a duration |
| Longer than the limit | Add moments; Argil guide cites 60 paragraphs | Split into several jobs |
| Avatar per video | avatarId on each moment | One resolved avatar per final video |
How should I split a script for each?
For Argil, split on sentence boundaries and count characters before you send, so no moment is rejected. Keep each moment a complete thought, because the moment is also the unit you can attach gestures, zoom or b-roll to.
For Sume, estimate seconds rather than characters. A short product pitch fits one job. A longer explainer becomes several jobs that you submit with distinct Idempotency-Key headers and join afterwards. Sume's Jobs and results page says a queued job is a normal accepted state, so submit the pieces together and poll or use a webhook rather than waiting on one at a time.
What does Sume not do here?
Sume does not give you a single avatar video longer than 60 seconds, and the docs state that one final video uses one resolved avatar. If you need a ten-minute talking presenter in one render, Argil's moment model is built for that length and Sume's is not. The Sume route fits clips: hooks, product pitches, short explainers, and per-recipient variants. See AI avatar video longer than 60 seconds for the join-it-yourself approach.
Sources
Related posts
More in Comparisons
- Argil subtitles Top/Middle/Bottom and sizes vs Sume caption design
Argil subtitles take a styleId, a position and a size. Sume burns captions with a named style plus design overrides. Compare the knobs before you pick.
- Argil video status IDLE to DONE vs Sume queued to completed
Argil videos move through IDLE, GENERATING_AUDIO, GENERATING_VIDEO, DONE or FAILED; Sume jobs go queued, processing, completed. How to map polling code.
- Argil VIDEO_GENERATION_SUCCESS webhook vs Sume job.completed payload
Argil sends four webhook events with videoUrl in data; Sume sends signed terminal job.completed, job.failed and job.canceled events. Handler differences.
- AssemblyAI Universal-3.5 Pro $0.21 per hour vs Sume STT
AssemblyAI lists Universal-3.5 Pro async at $0.21 an hour plus $0.02 for speaker labels. Sume STT is $0.60 an hour with no add-on line. What to compare.
Written by Sume