Muse Voice Transcribe 16 kHz mono PCM: ffmpeg vs Sume audio-detach
Muse Voice Transcribe wants mono 16-bit PCM at 24 or 16 kHz. On Sume, audio-detach with 16000 Hz and mono makes the STT input wav from a hosted video.

Meta's Muse Voice Transcribe guide says it supports mono 16-bit PCM audio at 24 kHz or 16 kHz and gives an ffmpeg command to convert anything else. On Sume, POST /v1/audio-detach with sample_rate: 16000 and channels: "mono" produces the 16 kHz mono wav that the docs call the STT shape.
Meta's input rule is from its developer guide; Sume's from the audio detach docs and the API reference, read 2026-10-01. Sume's STT is a separate request and does not stream.
What does Meta's guide ask for?
Mono 16-bit PCM at 24 kHz (engine-native) or 16 kHz. Its conversion line is ffmpeg -i input.mp3 -ac 1 -ar 24000 -sample_fmt s16 sample.wav. The 24 kHz option is Meta's and has no counterpart in the Sume audio-detach docs, which list only 16000, 44100 and 48000.
How do the same fields map to Sume?
Audio detach takes a Sume-hosted video, so the source must already be a media.sume.com artifact. The ffmpeg flags become request fields.
| Goal | Meta guide (ffmpeg) | Sume audio-detach field |
|---|---|---|
| Mono | -ac 1 | channels: "mono" (default source) |
| 16 kHz | -ar 16000 allowed | sample_rate: 16000 |
| 24 kHz | -ar 24000 | Not a listed value |
| 16-bit PCM | -sample_fmt s16 | format: "wav" (default, pcm_s16le) |
What does the request look like?
Send video_url, format, channels and sample_rate. Idempotency-Key is required, and the default mode is async; poll the job for the new audio_url.
curl -X POST https://api.sume.com/v1/audio-detach \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: detach-16k-001" \
-d '{
"video_url": "https://media.sume.com/artifacts/artf_demo/talk.mp4",
"format": "wav",
"channels": "mono",
"sample_rate": 16000
}'What does it cost and what are the caps?
Audio detach is $0.01 per job, with no provider inference, only worker ffmpeg. The source must be at most 1800 seconds and the output at most 900 seconds; past 900 seconds pass a range. See audio detach price.
Do I need this before Sume STT?
Not always. Sume STT takes a public HTTPS audio_url, and its transcript result keeps the 16 kHz mono wav it was made from. Detach is for when your source is a video and you want that wav yourself. The longer walkthrough is 16 kHz mono audio for speech-to-text.
Sources
Related posts
More in Developers
- Muse Voice Transcribe WebSocket streaming vs Sume STT job waits
Muse Voice Transcribe streams audio over one WebSocket and returns cumulative partials. Sume STT is a job: wait up to 30 seconds, then poll status_url.
- Nano Banana batch API: Gemini's 24 h batch vs Sume async jobs
Gemini's Batch API trades up to 24 hours of turnaround for higher rate limits. Sume has no batch tier for images: send async or webhook jobs per request.
- Nano Banana Pro 21:9: Gemini's ratio list vs Sume's catalog
Gemini's image docs list 21:9 among ten ratios. Sume's Nano Banana Pro and Nano Banana 2 catalogs include 21:9 too; GPT Image 2.5's list does not.
- Next.js dev MCP endpoint leak: keep your Sume API key out of source
CVE-2026-94486 let a malicious page read source snippets from next dev. Keep the Sume API key in an env var, server-side, and log request ids not headers.
Written by Sume