Argil audioUrl is limited to 40 seconds; what Sume takes instead
Argil's moment audioUrl is capped at 40 seconds. Sume's avatar talking-video takes script text or silence beats, not an audio file. What that means.
Argil lets you drive a moment with your own recorded audio through audioUrl, but the external audio is limited to 40 seconds. Sume's avatar talking-video route does not take an audio file at all in the documented contract: each spoken scene carries script text, and a silence scene carries only a duration.
The facts below come from Argil's Create a new Video page and Sume's Generate avatar video, both read on 2026-10-02.
What does Argil's audioUrl accept?
Each Argil moment takes either a transcript or an audioUrl, and the two are mutually exclusive. The page lists a 40-second limit on external audio. A moment can also carry an optional voice object, a gestureSlug, zoom keyframes (1 to 20 per moment) and broll.
So if you already have a recording, you can feed it in, but a long recording has to be cut into pieces of at most 40 seconds, one per moment.
What does the Sume avatar route take?
On POST /v1/avatar-1.0/talking-video you send avatar_handle, then exactly one of script or video_inputs. Inside video_inputs, a spoken scene uses voice.type: "text" with exactly one of script or input_text, and a non-speaking scene uses voice.type: "silence" with a required duration. The docs do not describe an audio-file field on this route.
Media fields elsewhere on the route, such as product_image and a photo scene, must be fetchable public HTTPS URLs; see Media inputs.
Which one fits which job?
The table summarizes the input choice.
| Need | Argil | Sume |
|---|---|---|
| Use a recording you already made | audioUrl, up to 40 seconds per moment | Not on the talking-video route as documented |
| Generate speech from text | transcript, up to 500 characters per moment | script or input_text per scene |
| Leave a gap with no speech | Not documented on the pages read | silence scene with duration |
| Total length | Several moments | 4-60 seconds estimated |
What should I do if I need my own recorded voice?
Check the other Sume routes before assuming it cannot be done. Sume has a separate audio-driven talking path covered in Audio to avatar AI, with its own limits, and a comparison of Argil's 50 MB upload limit against it. Treat those as different routes from the Avatar 1.0 talking-video endpoint, and read their pages before you build.
If your goal is only a consistent voice across clips, the simpler plan is to keep one avatar handle and write scripts, then check which voice your avatar speaks with.
Sources
Related posts
More in Comparisons
- Argil 500-character moment limit vs Sume's 4-60 second script window
Argil caps each moment transcript at 500 characters; Sume sizes a script by estimated seconds instead. How to split a long avatar script on either one.
- Argil subtitles Top/Middle/Bottom and sizes vs Sume caption design
Argil subtitles take a styleId, a position and a size. Sume burns captions with a named style plus design overrides. Compare the knobs before you pick.
- Argil video status IDLE to DONE vs Sume queued to completed
Argil videos move through IDLE, GENERATING_AUDIO, GENERATING_VIDEO, DONE or FAILED; Sume jobs go queued, processing, completed. How to map polling code.
- Argil VIDEO_GENERATION_SUCCESS webhook vs Sume job.completed payload
Argil sends four webhook events with videoUrl in data; Sume sends signed terminal job.completed, job.failed and job.canceled events. Handler differences.
Written by Sume