Argil audio upload 50 MB vs Sume Fabric's 10 MB audio URL
Argil takes mp3, wav and m4a uploads up to 50 MB. Sume's Fabric route takes a Sume-hosted audio URL up to 10 MB; H3 Max lip sync wants 5-14.8 seconds.

Argil's audio page says supported formats are mp3, wav and m4a with a maximum size of 50 MB. Sume's Fabric route takes an audio_url instead of an upload: it must be a public HTTPS URL on the Sume media host, typically a TTS segment, with a maximum of 10 MB. The H3 Max lip-sync route uses the same body with audio of 5-14.8 seconds.
Argil facts are from its audio page; Sume facts from the OpenAPI schema and Models, read 2026-10-01.
What are the audio-in limits side by side?
| Limit | Argil | Sume Fabric / H3 Max |
|---|---|---|
| Input style | File upload | audio_url string |
| Formats | mp3, wav, m4a | Not listed in the field description |
| Max size | 50 MB | 10 MB |
| Host | Your upload | Sume media host only; other hosts are rejected |
| Duration | Not stated on the page | H3 Max: 5-14.8 seconds |
What does the Sume host rule mean in practice?
Non-Sume hosts are rejected, so a file on your own server or bucket will not work as-is. The field description points at a TTS segment as the typical source: generate speech with Sume TTS and pass the resulting media URL. The route itself is POST /v1/veed/fabric-1.0, a talking still plus audio clip, and the lip-sync route is POST /v1/minimax/h3-max/lip-sync.
What does Argil do with the audio?
After upload, the page says Argil transcribes the audio and lets you transform your voice while preserving emotions and tone. Sume's Fabric request does not describe a transcription step in the sources read; it takes the audio as given.
What should I do if my audio is too big or too long?
Split it into segments that fit, render each, and join the clips afterwards; for a similar limit comparison see HeyGen audio assets vs Fabric. Check the size before you submit, since the cap is on the file, and poll the job as usual: status_url until terminal, then result_url.
Sources
Related posts
More in Developers
- AssemblyAI disfluencies: true (English only) vs Sume STT words[]
AssemblyAI's disfluencies: true raised filler-word recall from 49.7% to 76.3%, English only. Sume STT has no such flag, so check words[] on your clip.
- AssemblyAI format_text false: 'one hundred' vs 100 and Sume STT
AssemblyAI's format_text: false returns spoken form, 'one hundred' not 100. Sume STT has no formatting toggle; its request is a closed schema.
- AssemblyAI speech models: unpinned routing and Sume's fixed STT id
AssemblyAI will route unpinned async requests to Universal-3.5 Pro; pin speech_models to stay put. What changes, and why Sume STT has no model to pin.
- AssemblyAI Sync API: one call, and Sume STT sync mode
AssemblyAI's Sync API returns a short-clip transcript in one POST. Sume STT mode sync waits up to 30 seconds, then returns the job id to poll if not done.
Written by Sume