TTS raw PCM output: container raw needs sample_rate and encoding
Sume TTS can return mp3, wav or headerless raw audio. Raw is the only container where sample_rate and encoding are required, and the fields are strict.

Set output_format to {"container": "raw", "sample_rate": 16000, "encoding": "pcm_s16le"} on a Sume TTS request. Raw output has no file header, so both sample_rate and encoding are required and have no default. The mp3 and wav containers have defaults (44100 Hz), which is why people who switch to raw usually hit a 400 on the first try.
The three containers
Sume TTS 1.0 and the TTS Router share one output schema. It is a discriminated union on container, and each branch is strict, meaning an extra field is rejected rather than ignored.
| Container | sample_rate | Other field | Defaults |
|---|---|---|---|
| mp3 | optional | bit_rate: 32000, 64000, 96000, 128000 or 192000 | 44100 Hz, 128000 |
| wav | optional | encoding | 44100 Hz, pcm_s16le |
| raw | required | encoding required | none |
Sample rates and encodings
The accepted sample rates are 8000, 16000, 22050, 24000, 44100 and 48000. The accepted encodings are pcm_f32le, pcm_s16le, pcm_mulaw and pcm_alaw. The same lists apply to wav and raw, so you can ask for the same audio in either container; only the header differs.
A speech pipeline that expects 16 kHz 16-bit samples gets that with raw plus pcm_s16le. A phone system that wants companded audio at 8 kHz gets it with pcm_mulaw or pcm_alaw. Both avoid a conversion step on your side.
A request
Set VOICE_ID to a voice id from your workspace; an id that is not a UUID or a voi_ id with 32 hex characters is rejected with 400 invalid_voice_id.
curl -X POST https://api.sume.com/v1/tts-1.0/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: tts-raw-001" \
-d '{
"transcript": "Your order has shipped.",
"voice": { "id": "'"$VOICE_ID"'" },
"language": "en",
"output_format": { "container": "raw", "sample_rate": 16000, "encoding": "pcm_s16le" }
}'What to check in the result
The result echoes output_format, so confirm it is what you asked for before you treat the bytes as PCM. If you send bit_rate with a wav or raw container, or omit sample_rate on raw, the strict schema returns a validation error naming the field.
Raw files have no duration header. Read the duration from the job result, not from the file, and keep the sample rate and encoding with the file name, since nothing inside the file says which they are.
When raw is the wrong choice
If the file goes into a Sume Timeline render or a lip-sync job, use wav or mp3: those jobs read a normal audio file. Raw is for your own downstream code. For joins and cuts, wav keeps timing exact. Sume bills the same $0.0475 per 1,000 characters whichever container you choose. Cartesia's Sonic 3.6 page is the vendor reference for the model behind these voices.
A small checklist
Before you wire raw output into code, confirm four things: the container is raw, the sample rate is in the allowed list, the encoding is one of the four, and no mp3-only field such as bit_rate is present.
Save one raw file and open it as headerless audio in an editor with the same settings. If it sounds right, your pipeline will too.
Sources
Related posts
More in Developers
- TTS Router 400 unknown model: which Sonic ids does Sume accept?
POST /v1/tts-router/generate needs a catalog model id; an unknown one returns 400 with catalog_url. Seed ids: sonic-3.6, 3.5, 3, latest, preview. Curl inside.
- Korean TTS segment text has no spaces, unless a digit is in it
Sume's TTS segment text joins tokens with spaces only if one has a Latin letter or digit; else with nothing. Use segments for timing, your script for text.
- TTS sentence slices: mp3 gives timings only, wav gives audio_urls
Sume TTS segmentation returns sentence timings for any container, but slice audio_urls only with wav or raw. Request shape, the 70 ms rule and when to pick wav.
- Pipes and @{} markers in a Sume TTS transcript: stripped, never spoken
Sume strips || cue breaks, @{...} markers and the display side of <display|spoken> before the voice reads; an empty result returns 400 transcript_no_speech.
Written by Sume