Grok TTS codecs and sample rates vs Sume TTS output_format
xAI TTS and Sume TTS offer the same six sample rates and mp3 bit-rate range. They differ on defaults (24 kHz vs 44.1 kHz) and Sume adds raw and float PCM.

Do Grok TTS and Sume TTS output the same audio formats?
Nearly. The xAI text-to-speech docs list MP3, WAV, PCM, mu-law and A-law, and six sample rates. Sume TTS 1.0 exposes the same six sample rates and a matching mp3 bit-rate range through its output_format object. The visible differences are the default sample rate and how encodings are named.
That makes a Grok-to-Sume port mostly a matter of renaming fields, with one trap: if you never set a sample rate, you get 24 kHz on xAI and 44.1 kHz on Sume.
What does each side list?
Values below come from the xAI docs page and the Sume OpenAPI contract for POST /v1/tts-1.0/generate. Neither side lists Opus or FLAC in the parts read.
| Option | xAI TTS | Sume TTS 1.0 |
|---|---|---|
| Containers or codecs | MP3, WAV, PCM, mu-law, A-law | mp3, wav, raw |
| PCM encodings | PCM, mu-law, A-law | pcm_f32le, pcm_s16le, pcm_mulaw, pcm_alaw |
| Sample rates (Hz) | 8000, 16000, 22050, 24000, 44100, 48000 | 8000, 16000, 22050, 24000, 44100, 48000 |
| Default sample rate | 24000 | 44100 |
| MP3 bit rates | 32000 to 192000, default 128000 | 32000, 64000, 96000, 128000, 192000; default 128000 |
How do you ask Sume for phone-quality audio?
For a telephony test you want 8 kHz mu-law in a WAV container. Set the container, sample rate and encoding together; the contract says encoding applies to wav and raw containers.
curl -X POST https://api.sume.com/v1/tts-1.0/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: phone-prompt-001" \
-d '{
"transcript": "Thanks for calling. Your order has shipped.",
"avatar_handle": "narrator",
"output_format": {
"container": "wav",
"sample_rate": 8000,
"encoding": "pcm_mulaw"
}
}'Which format should you pick for which job?
Use mp3 at 44.1 kHz and 128 kbps, Sume's default, for audio people will download or embed. Use wav with pcm_s16le at 44,100 or 48,000 Hz when the file will be edited, joined or fed to a lip-sync step, because the contract recommends wav/pcm_s16le/44100 explicitly when feeding an avatar mux.
Sentence slices follow the same rule. segmentation with emit_audio produces per-sentence audio files only for wav or raw containers; with mp3 you still get timings but no slice files. If you want word timings for captions, ask for timestamps.words as well.
pcm_f32le is the Sume-only option here: 32-bit float PCM suits a DSP chain that wants headroom, but the files are large, so convert to 16-bit for delivery.
What should you verify after porting?
Play one sample at each rate you intend to ship; do not assume equal loudness between engines. Sume's volume control is generation_config.volume from 0.5 to 2.0, so level-match there rather than re-encoding.
Check duration too. A TTS job whose synthesized audio exceeds 1,200 seconds fails with tts_duration_exceeded, whichever format you chose. The sample rate post has the long-form workflow.
Sources
Related posts
More in Comparisons
- Grok TTS language "auto" vs Sume's explicit language field
xAI TTS can auto-detect the language of your text. Sume TTS cannot: omit language and Spanish is read as English, except Korean and Japanese.
- Grok TTS speech tags like laugh and whisper vs Sume's emotion field
xAI lets you write [laugh] and wrap text in whisper or singing tags. Sume TTS documents no tag grammar, only emotion, speed and volume in generation_config.
- H3 Max Recast vs Genjutsu: which person swap to call on Sume
Sume lists two person-swap video rows. Recast: 1-4 people in a 5-30 s clip at 768p or 1080p. Genjutsu: 1-8 images at 480p or 720p. How to choose.
- Hedra 402 INSUFFICIENT_BALANCE vs Sume insufficient_credits
Hedra returns 402 INSUFFICIENT_BALANCE until you add funds; Sume returns 402 insufficient_credits. How each wallet check behaves and how to preflight a job.
Written by Sume