Muse Voice Transcribe 16 kHz mono PCM: ffmpeg vs Sume audio-detach

Muse Voice Transcribe wants mono 16-bit PCM at 24 or 16 kHz. On Sume, audio-detach with 16000 Hz and mono makes the STT input wav from a hosted video.

4 min readSume
All posts

Meta's Muse Voice Transcribe guide says it supports mono 16-bit PCM audio at 24 kHz or 16 kHz and gives an ffmpeg command to convert anything else. On Sume, POST /v1/audio-detach with sample_rate: 16000 and channels: "mono" produces the 16 kHz mono wav that the docs call the STT shape.

Meta's input rule is from its developer guide; Sume's from the audio detach docs and the API reference, read 2026-10-01. Sume's STT is a separate request and does not stream.

What does Meta's guide ask for?

Mono 16-bit PCM at 24 kHz (engine-native) or 16 kHz. Its conversion line is ffmpeg -i input.mp3 -ac 1 -ar 24000 -sample_fmt s16 sample.wav. The 24 kHz option is Meta's and has no counterpart in the Sume audio-detach docs, which list only 16000, 44100 and 48000.

How do the same fields map to Sume?

Audio detach takes a Sume-hosted video, so the source must already be a media.sume.com artifact. The ffmpeg flags become request fields.

ffmpeg flags and Sume audio-detach fields, Meta guide and Sume docs, read 2026-10-01.
GoalMeta guide (ffmpeg)Sume audio-detach field
Mono-ac 1channels: "mono" (default source)
16 kHz-ar 16000 allowedsample_rate: 16000
24 kHz-ar 24000Not a listed value
16-bit PCM-sample_fmt s16format: "wav" (default, pcm_s16le)

What does the request look like?

Send video_url, format, channels and sample_rate. Idempotency-Key is required, and the default mode is async; poll the job for the new audio_url.

curl -X POST https://api.sume.com/v1/audio-detach \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: detach-16k-001" \
  -d '{
    "video_url": "https://media.sume.com/artifacts/artf_demo/talk.mp4",
    "format": "wav",
    "channels": "mono",
    "sample_rate": 16000
  }'

What does it cost and what are the caps?

Audio detach is $0.01 per job, with no provider inference, only worker ffmpeg. The source must be at most 1800 seconds and the output at most 900 seconds; past 900 seconds pass a range. See audio detach price.

Do I need this before Sume STT?

Not always. Sume STT takes a public HTTPS audio_url, and its transcript result keeps the 16 kHz mono wav it was made from. Detach is for when your source is a video and you want that wav yourself. The longer walkthrough is 16 kHz mono audio for speech-to-text.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume