TTS sentence slices have no audio_url: emit_audio needs wav or raw

With segmentation on, Sume TTS only returns a sample-exact audio_url per sentence for wav or raw output. For mp3 you get timings and a warning instead.

4 min readSume
All posts

If your Sume TTS sentence segments come back without an audio_url, the job was asked for mp3. segmentation.emit_audio only slices audio when output_format.container is wav or raw; for mp3 the job still returns each sentence's timing, leaves out the per-segment audio_url, and adds a warning. Resubmit with output_format: { "container": "wav", "encoding": "pcm_s16le", "sample_rate": 44100 } and keep timestamps.words set to true.

What are the exact rules?

The OpenAPI description says segmentation requires timestamps.words: true, returns gapless segments[] where each segment's end equals the next one's start, and that emit_audio defaults to true. It also says: when emit_audio is true and the container is wav or raw, each segment includes a sample-exact audio_url; for mp3, timings are returned without it.

In the worker code the warning text is: "segmentation.emit_audio requires wav or raw output_format.container; segment audio_url omitted for mp3." It appears in the result's warnings array, so a job that looks fine can still be missing the files.

Behavior of segmentation.mode sentence by container, from the Sume OpenAPI and TTS worker code, read 2026-10-02.
ContainerSegment timingsSegment audio_urlWarning
wavYesYes, sample-exactNo
rawYesYesNo
mp3 (the default)YesNoYes, in warnings

Why does the default bite?

TTS defaults to mp3 at 44100 Hz and 128 kbps. The request description tells you to pass wav with pcm_s16le at 44100 explicitly when the audio feeds an avatar mux. Segment slices are the same case: you want clean cut points, and a wav file gives you them.

Cartesia's changelog (read 2026-10-02) says for Sonic 3.6 that timestamps and codec behave as they do on Sonic 3.5, so nothing about the model changes this rule.

What does a working request look like?

Request wav, ask for word timings, then turn on sentence segmentation. Each entry in segments[] has index, text, start, end and duration_seconds, plus audio_url when slices were made.

curl -X POST https://api.sume.com/v1/tts-1.0/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: seg-wav-001" \
  -d '{
    "transcript": "Welcome back. Today we cover three updates.",
    "avatar_handle": "@your_avatar",
    "language": "en",
    "output_format": { "container": "wav", "encoding": "pcm_s16le", "sample_rate": 44100 },
    "timestamps": { "words": true },
    "segmentation": { "mode": "sentence", "emit_audio": true }
  }'

Can I convert to mp3 later?

Yes, after the cuts are made. Join or slice with Timeline audio, which works on Sume-hosted audio in the sample domain, and hand out mp3 only as the last step. Keep the wav slices as your masters.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume