TTS raw PCM output: container raw needs sample_rate and encoding

Sume TTS can return mp3, wav or headerless raw audio. Raw is the only container where sample_rate and encoding are required, and the fields are strict.

5 min readSume
All posts

Set output_format to {"container": "raw", "sample_rate": 16000, "encoding": "pcm_s16le"} on a Sume TTS request. Raw output has no file header, so both sample_rate and encoding are required and have no default. The mp3 and wav containers have defaults (44100 Hz), which is why people who switch to raw usually hit a 400 on the first try.

The three containers

Sume TTS 1.0 and the TTS Router share one output schema. It is a discriminated union on container, and each branch is strict, meaning an extra field is rejected rather than ignored.

TTS output_format fields from apps/api/src/schemas.ts (read 2026-10-05)
Containersample_rateOther fieldDefaults
mp3optionalbit_rate: 32000, 64000, 96000, 128000 or 19200044100 Hz, 128000
wavoptionalencoding44100 Hz, pcm_s16le
rawrequiredencoding requirednone

Sample rates and encodings

The accepted sample rates are 8000, 16000, 22050, 24000, 44100 and 48000. The accepted encodings are pcm_f32le, pcm_s16le, pcm_mulaw and pcm_alaw. The same lists apply to wav and raw, so you can ask for the same audio in either container; only the header differs.

A speech pipeline that expects 16 kHz 16-bit samples gets that with raw plus pcm_s16le. A phone system that wants companded audio at 8 kHz gets it with pcm_mulaw or pcm_alaw. Both avoid a conversion step on your side.

A request

Set VOICE_ID to a voice id from your workspace; an id that is not a UUID or a voi_ id with 32 hex characters is rejected with 400 invalid_voice_id.

curl -X POST https://api.sume.com/v1/tts-1.0/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: tts-raw-001" \
  -d '{
    "transcript": "Your order has shipped.",
    "voice": { "id": "'"$VOICE_ID"'" },
    "language": "en",
    "output_format": { "container": "raw", "sample_rate": 16000, "encoding": "pcm_s16le" }
  }'

What to check in the result

The result echoes output_format, so confirm it is what you asked for before you treat the bytes as PCM. If you send bit_rate with a wav or raw container, or omit sample_rate on raw, the strict schema returns a validation error naming the field.

Raw files have no duration header. Read the duration from the job result, not from the file, and keep the sample rate and encoding with the file name, since nothing inside the file says which they are.

When raw is the wrong choice

If the file goes into a Sume Timeline render or a lip-sync job, use wav or mp3: those jobs read a normal audio file. Raw is for your own downstream code. For joins and cuts, wav keeps timing exact. Sume bills the same $0.0475 per 1,000 characters whichever container you choose. Cartesia's Sonic 3.6 page is the vendor reference for the model behind these voices.

A small checklist

Before you wire raw output into code, confirm four things: the container is raw, the sample rate is in the allowed list, the encoding is one of the four, and no mp3-only field such as bit_rate is present.

Save one raw file and open it as headerless audio in an editor with the same settings. If it sounds right, your pipeline will too.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume