Azure fast transcription: 500 MB, under 5 h, vs Sume's 900 s detach
Azure fast transcription takes audio under 500 MB and under 5 hours. Sume audio detach outputs at most 900 s, so longer audio needs ranges.

Azure's fast transcription API accepts audio files under 500 MB and under 5 hours. Sume's audio detach writes at most 900 seconds of audio per job, from a source video of up to 1800 seconds, so detaching is a way to prepare a track for Azure only when the part you need fits inside those caps. For anything longer, use range to take it in pieces.
Azure's limits are from its fast transcription page, read 2026-10-03. Sume's are from audio detach.
The two ceilings
The limits are of different kinds: Azure's are about the file you upload, and Sume's are about the audio a job may produce.
| Limit | Azure fast transcription | Sume audio detach |
|---|---|---|
| Maximum file size | Under 500 MB | Not stated on the detach page |
| Maximum duration | Less than 5 hours | Source 1800 s, output 900 s |
| Whole track longer than the cap | Not allowed | Needs a range |
Taking a long recording in pieces
The 5 hour limit is about 18000 seconds, while Sume's source cap is 1800 seconds. That means Sume's detach will not read a 2 hour video at all, because the worker refuses a source longer than 1800 s with source_duration_exceeded. If your video is 30 minutes or less, you can take it in ranges of up to 900 s each.
Each range is its own $0.01 job. Two ranges cover a 30 minute video: { start: 0, end: 900 } and { start: 900, end: 1800 }. Name your files so the order is clear when you send them on.
The audio shape to ask for
For speech use channels: "mono" and sample_rate: 16000, which Sume's docs call the STT shape, in the default wav format. Azure's page says nothing about which of these you must send, so check Azure's supported formats before choosing mp3.
Because wav here is pcm_s16le, one minute of 16 kHz mono audio is 16000 samples a second at 2 bytes each, about 1.92 MB, by arithmetic. A 900 second range is about 29 MB, well under 500 MB.
- Check
probe.has_audiofirst with video inspect andframes: false. - A source with no audio fails
detach_source_has_no_audio. - The video file is untouched.
When Sume can transcribe for you
Video inspect has an optional transcript through Sume STT 1.0 at $0.01 per audio minute. If you only need text, that route skips the hand-off. If you need Azure features or a specific region, use detach and send the file there.
Naming and ordering the pieces
When you detach a long video in ranges, keep a list of each range's start and end next to the returned audio_url. After transcription, add each range's start to that piece's word times, so the timings line up with the original video.
Without that offset, every piece restarts at zero and the transcript is hard to align.
Which service holds the long file
The two limits tell you where a long recording should live. A 3 hour interview is within Azure's fast transcription limits as audio, but Sume's detach refuses a source video longer than 1800 seconds, so it cannot prepare that file.
So for long recordings, extract the audio before it reaches Sume, using a tool outside it, and send that straight to Azure.
Sources
Related posts
More in Media tools
- Batch trim clips from a spreadsheet of start and end times
Read start and end columns from a CSV and trim a long video into Shorts: one Video trim call per row at $0.02, with idempotency keys.
- Black bars on a YouTube Short: YouTube says no, so fill the frame
YouTube says uploads should never include letterbox or pillarbox bars, and Shorts take square or vertical files. Reframe 16:9 clips with Sume Timeline fit.
- Burned-in captions: ALL CAPS or as written? What each Sume style does
Sume slam, punch and tiktok-green upper-case every Latin word at render time; Hangul-identity styles burn text as written. What script_text and cues change.
- Can a 16:9 video be an Instagram Reel? The range says yes
Instagram accepts Reels from 1.91:1 to 9:16, so 16:9 (1.78) is inside the range. How to render a 1920x1080 cut or a vertical one from the same Sume clips.
Written by Sume