Grok TTS codecs and sample rates vs Sume TTS output_format

xAI TTS and Sume TTS offer the same six sample rates and mp3 bit-rate range. They differ on defaults (24 kHz vs 44.1 kHz) and Sume adds raw and float PCM.

5 min readSume
All posts

Do Grok TTS and Sume TTS output the same audio formats?

Nearly. The xAI text-to-speech docs list MP3, WAV, PCM, mu-law and A-law, and six sample rates. Sume TTS 1.0 exposes the same six sample rates and a matching mp3 bit-rate range through its output_format object. The visible differences are the default sample rate and how encodings are named.

That makes a Grok-to-Sume port mostly a matter of renaming fields, with one trap: if you never set a sample rate, you get 24 kHz on xAI and 44.1 kHz on Sume.

What does each side list?

Values below come from the xAI docs page and the Sume OpenAPI contract for POST /v1/tts-1.0/generate. Neither side lists Opus or FLAC in the parts read.

Output options, xAI docs and Sume OpenAPI (read 2026-10-02)
OptionxAI TTSSume TTS 1.0
Containers or codecsMP3, WAV, PCM, mu-law, A-lawmp3, wav, raw
PCM encodingsPCM, mu-law, A-lawpcm_f32le, pcm_s16le, pcm_mulaw, pcm_alaw
Sample rates (Hz)8000, 16000, 22050, 24000, 44100, 480008000, 16000, 22050, 24000, 44100, 48000
Default sample rate2400044100
MP3 bit rates32000 to 192000, default 12800032000, 64000, 96000, 128000, 192000; default 128000

How do you ask Sume for phone-quality audio?

For a telephony test you want 8 kHz mu-law in a WAV container. Set the container, sample rate and encoding together; the contract says encoding applies to wav and raw containers.

curl -X POST https://api.sume.com/v1/tts-1.0/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: phone-prompt-001" \
  -d '{
    "transcript": "Thanks for calling. Your order has shipped.",
    "avatar_handle": "narrator",
    "output_format": {
      "container": "wav",
      "sample_rate": 8000,
      "encoding": "pcm_mulaw"
    }
  }'

Which format should you pick for which job?

Use mp3 at 44.1 kHz and 128 kbps, Sume's default, for audio people will download or embed. Use wav with pcm_s16le at 44,100 or 48,000 Hz when the file will be edited, joined or fed to a lip-sync step, because the contract recommends wav/pcm_s16le/44100 explicitly when feeding an avatar mux.

Sentence slices follow the same rule. segmentation with emit_audio produces per-sentence audio files only for wav or raw containers; with mp3 you still get timings but no slice files. If you want word timings for captions, ask for timestamps.words as well.

pcm_f32le is the Sume-only option here: 32-bit float PCM suits a DSP chain that wants headroom, but the files are large, so convert to 16-bit for delivery.

What should you verify after porting?

Play one sample at each rate you intend to ship; do not assume equal loudness between engines. Sume's volume control is generation_config.volume from 0.5 to 2.0, so level-match there rather than re-encoding.

Check duration too. A TTS job whose synthesized audio exceeds 1,200 seconds fails with tts_duration_exceeded, whichever format you chose. The sample rate post has the long-form workflow.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume