Azure batch transcription wants WAV PCM or FLAC: Sume detach's default
Azure advises lossless WAV (PCM) or FLAC for best transcription quality. Sume audio detach defaults to sample-exact pcm_s16le wav, so the default fits.

Azure's batch transcription page lists many accepted formats, and says that for the best transcription quality you should use lossless formats such as WAV with PCM encoding and FLAC. Sume's audio detach returns a sample-exact pcm_s16le wav by default, so you get the recommended shape without setting anything; mp3 at 128 kbps is the other option, and the lossy one.
The Azure advice is from its batch transcription audio data page, read 2026-10-03. The Sume defaults are from audio detach.
What Azure accepts
The page lists a long set of containers and codecs and then says the service may accept more, because it integrates GStreamer.
| Group | Formats named on the page |
|---|---|
| Lossless, recommended | WAV (PCM encoding), FLAC |
| Compressed | MP3, OPUS/OGG, AAC, WMA, AMR, WebM, SPEEX |
| Telephony | ALAW in WAV container, MULAW in WAV container |
Where mp3 fits
The page names mp3 as accepted but recommends lossless for quality. Sume offers mp3 at 128 kbps, which is smaller to move around. If transcription quality matters more than upload size, keep the default wav. If the file is large and the audio is clean, mp3 may be good enough, and a test on your own audio will show it.
Sume's timeline audio docs add a second reason to stay with wav: mp3 re-adds priming padding at every edge, which matters if the file will be joined or split again.
The detach request for Azure
Import the video with POST /v1/media-imports, then post the media.sume.com URL to /v1/audio-detach. The defaults give wav with the source channel layout and sample rate. For speech, set channels to mono and sample_rate to 16000 as Sume's docs recommend for speech-to-text.
Azure's page does not state a preferred sample rate, so test before you commit to a rate for a large batch.
format:wavdefault, ormp3.range: a start and optional end in seconds.- Cap: 900 s of output per job, 1800 s of source.
Pointing Azure at the file
Azure's batch service reads from a public URI that needs no authentication, or from Blob storage through a SAS URL. A media.sume.com audio URL is not an Azure Blob address, so check that your account can fetch it, or copy the file into your own storage first. The page says the service does not support URIs that require authentication.
What Sume does not offer
Sume's detach has no FLAC option: the docs list wav and mp3 only. If you want FLAC for Azure, convert the wav yourself after detaching. Because the wav is lossless, converting it to FLAC loses nothing.
Azure also says it may accept more formats than listed, but a list is not a promise, so keep to the named lossless formats when quality matters.
Sample rate and channels
Azure's page does not tell you which sample rate to use, so there is no figure to cite. Sume's docs say 16000 Hz mono is the speech-to-text shape and that omitting sample_rate inherits the source. Inheriting is the safe choice if you are unsure, because it avoids a conversion step on your side.
If your recordings are stereo with one speaker per channel, keep channels at source rather than forcing mono, so you do not mix the voices together.
Sources
Related posts
More in Media tools
- Azure fast transcription: 500 MB, under 5 h, vs Sume's 900 s detach
Azure fast transcription takes audio under 500 MB and under 5 hours. Sume audio detach outputs at most 900 s, so longer audio needs ranges.
- Batch trim clips from a spreadsheet of start and end times
Read start and end columns from a CSV and trim a long video into Shorts: one Video trim call per row at $0.02, with idempotency keys.
- Black bars on a YouTube Short: YouTube says no, so fill the frame
YouTube says uploads should never include letterbox or pillarbox bars, and Shorts take square or vertical files. Reframe 16:9 clips with Sume Timeline fit.
- Burned-in captions: ALL CAPS or as written? What each Sume style does
Sume slam, punch and tiktok-green upper-case every Latin word at render time; Hangul-identity styles burn text as written. What script_text and cues change.
Written by Sume