TTS sentence slices have no audio_url: emit_audio needs wav or raw
With segmentation on, Sume TTS only returns a sample-exact audio_url per sentence for wav or raw output. For mp3 you get timings and a warning instead.

If your Sume TTS sentence segments come back without an audio_url, the job was asked for mp3. segmentation.emit_audio only slices audio when output_format.container is wav or raw; for mp3 the job still returns each sentence's timing, leaves out the per-segment audio_url, and adds a warning. Resubmit with output_format: { "container": "wav", "encoding": "pcm_s16le", "sample_rate": 44100 } and keep timestamps.words set to true.
What are the exact rules?
The OpenAPI description says segmentation requires timestamps.words: true, returns gapless segments[] where each segment's end equals the next one's start, and that emit_audio defaults to true. It also says: when emit_audio is true and the container is wav or raw, each segment includes a sample-exact audio_url; for mp3, timings are returned without it.
In the worker code the warning text is: "segmentation.emit_audio requires wav or raw output_format.container; segment audio_url omitted for mp3." It appears in the result's warnings array, so a job that looks fine can still be missing the files.
| Container | Segment timings | Segment audio_url | Warning |
|---|---|---|---|
wav | Yes | Yes, sample-exact | No |
raw | Yes | Yes | No |
mp3 (the default) | Yes | No | Yes, in warnings |
Why does the default bite?
TTS defaults to mp3 at 44100 Hz and 128 kbps. The request description tells you to pass wav with pcm_s16le at 44100 explicitly when the audio feeds an avatar mux. Segment slices are the same case: you want clean cut points, and a wav file gives you them.
Cartesia's changelog (read 2026-10-02) says for Sonic 3.6 that timestamps and codec behave as they do on Sonic 3.5, so nothing about the model changes this rule.
What does a working request look like?
Request wav, ask for word timings, then turn on sentence segmentation. Each entry in segments[] has index, text, start, end and duration_seconds, plus audio_url when slices were made.
curl -X POST https://api.sume.com/v1/tts-1.0/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: seg-wav-001" \
-d '{
"transcript": "Welcome back. Today we cover three updates.",
"avatar_handle": "@your_avatar",
"language": "en",
"output_format": { "container": "wav", "encoding": "pcm_s16le", "sample_rate": 44100 },
"timestamps": { "words": true },
"segmentation": { "mode": "sentence", "emit_audio": true }
}'Can I convert to mp3 later?
Yes, after the cuts are made. Join or slice with Timeline audio, which works on Sume-hosted audio in the sample domain, and hand out mp3 only as the last step. Keep the wav slices as your masters.
Sources
Related posts
More in Developers
- TTS speed: slow, normal, fast or generation_config.speed 0.6 to 1.5?
Sume TTS marks the slow, normal and fast speed enum deprecated. Send generation_config.speed from 0.6 to 1.5, plus volume 0.5 to 2 and an emotion string.
- TTS call returned processing in sync mode: poll, do not resubmit
Sume TTS in sync mode waits at most 30 seconds. If the job is not terminal, poll status_url, and retry a submit only with the same Idempotency-Key.
- Test call audio for voice agents: Sume TTS at 8 kHz mu-law
Generate repeatable phone-quality test utterances for a voice agent with Sume TTS output_format: 8000 Hz, pcm_mulaw. Fields, limits and a runnable script.
- TTS voice.id: a UUID or a voi_ library id? What Sume accepts
Sume TTS voice.id takes a voice UUID or a voi_ library id. Any other shape fails with 400 invalid_voice_id before a job is queued or credits are reserved.
Written by Sume