Azure batch transcription wants WAV PCM or FLAC: Sume detach's default

Azure advises lossless WAV (PCM) or FLAC for best transcription quality. Sume audio detach defaults to sample-exact pcm_s16le wav, so the default fits.

5 min readSume
All posts

Azure's batch transcription page lists many accepted formats, and says that for the best transcription quality you should use lossless formats such as WAV with PCM encoding and FLAC. Sume's audio detach returns a sample-exact pcm_s16le wav by default, so you get the recommended shape without setting anything; mp3 at 128 kbps is the other option, and the lossy one.

The Azure advice is from its batch transcription audio data page, read 2026-10-03. The Sume defaults are from audio detach.

What Azure accepts

The page lists a long set of containers and codecs and then says the service may accept more, because it integrates GStreamer.

Azure batch input formats (read 2026-10-03)
GroupFormats named on the page
Lossless, recommendedWAV (PCM encoding), FLAC
CompressedMP3, OPUS/OGG, AAC, WMA, AMR, WebM, SPEEX
TelephonyALAW in WAV container, MULAW in WAV container

Where mp3 fits

The page names mp3 as accepted but recommends lossless for quality. Sume offers mp3 at 128 kbps, which is smaller to move around. If transcription quality matters more than upload size, keep the default wav. If the file is large and the audio is clean, mp3 may be good enough, and a test on your own audio will show it.

Sume's timeline audio docs add a second reason to stay with wav: mp3 re-adds priming padding at every edge, which matters if the file will be joined or split again.

The detach request for Azure

Import the video with POST /v1/media-imports, then post the media.sume.com URL to /v1/audio-detach. The defaults give wav with the source channel layout and sample rate. For speech, set channels to mono and sample_rate to 16000 as Sume's docs recommend for speech-to-text.

Azure's page does not state a preferred sample rate, so test before you commit to a rate for a large batch.

  • format: wav default, or mp3.
  • range: a start and optional end in seconds.
  • Cap: 900 s of output per job, 1800 s of source.

Pointing Azure at the file

Azure's batch service reads from a public URI that needs no authentication, or from Blob storage through a SAS URL. A media.sume.com audio URL is not an Azure Blob address, so check that your account can fetch it, or copy the file into your own storage first. The page says the service does not support URIs that require authentication.

What Sume does not offer

Sume's detach has no FLAC option: the docs list wav and mp3 only. If you want FLAC for Azure, convert the wav yourself after detaching. Because the wav is lossless, converting it to FLAC loses nothing.

Azure also says it may accept more formats than listed, but a list is not a promise, so keep to the named lossless formats when quality matters.

Sample rate and channels

Azure's page does not tell you which sample rate to use, so there is no figure to cite. Sume's docs say 16000 Hz mono is the speech-to-text shape and that omitting sample_rate inherits the source. Inheriting is the safe choice if you are unsure, because it avoids a conversion step on your side.

If your recordings are stereo with one speaker per channel, keep channels at source rather than forcing mono, so you do not mix the voices together.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume