AssemblyAI Sync API file limit vs Sume STT duration_seconds

AssemblyAI's launch post frames its Sync API for short clips. Sume STT 1.0 takes duration_seconds from 1 to 600 and reserves one minute if you omit it.

4 min readSume
All posts

AssemblyAI's launch post describes its Sync API as built for short clips and says the clip is decoded to a single 16 kHz audio array with no chunking; the post text I read states no byte or minute cap, so check AssemblyAI's own docs for the exact file limit. Sume's STT 1.0 numbers are explicit: duration_seconds takes 1 to 600, and omitting it reserves 1 minute.

AssemblyAI facts are from its launch post; Sume facts from the OpenAPI schema and docs, read 2026-10-01.

What does AssemblyAI say about clip length?

The post lists the Sync API for "short clips: dictation, voice agents with turn detection" and sends "long audio, batch, latency-tolerant" work to the Async API. It does not give a number in the part I read, so I do not quote one.

What are Sume's length numbers?

Length limits that apply to Sume transcription, read 2026-10-01.
SettingValueSource
duration_seconds1 to 600; improves the usage reservationSTT schema
Omitted duration_secondsReserves 1 minuteSTT schema
Maximum hint10 minutesSTT schema
Audio Detach outputAt most 900 sAudio Detach docs
Rate$0.01 per audio minuteRate card

What happens if I leave duration_seconds out?

The reservation is one minute. For a 5-minute recording, send duration_seconds: 300 so the reservation matches the audio. The schema says the field improves the usage reservation; it is a hint for billing, not a cut-off you set on the audio.

What about audio longer than 10 minutes?

Split it into parts under 600 seconds and submit each as its own job, then join the word lists by offset. If the source is a video, Audio Detach can extract a 16 kHz mono track first; see audio detach for speech-to-text. Its own output limit is 900 s, so long files need a range there too.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume