Sume TTS sentence slices need wav or raw output, not mp3

Sume TTS returns per-sentence audio_url slices only for wav or raw output. With mp3 you get timings but no slices. The request that gets clips.

4 min readSume
All posts

If you ask Sume TTS for sentence segmentation and leave the output at the default, you get timings but no clips. Per-segment audio_url slices exist only when output_format.container is wav or raw. The default is mp3 at 44,100 Hz and 128 kbps, so set the container explicitly when you want one file per sentence.

The rule in the schema

The OpenAPI description for segmentation says it requires timestamps.words: true, returns gapless segments[] (each end equals the next start), and applies a 70 ms post-word boundary by default. emit_audio defaults to true, but "when true and output_format.container is wav/raw, each segment includes a sample-exact audio_url. For mp3, timings are returned without segment audio_url."

That is a deliberate split: lossy mp3 frames do not cut sample-exactly, so Sume only slices PCM.

What a segmented TTS job returns by container - from the Sume OpenAPI schema (read 2026-10-07)
ContainerSegment timingsPer-segment audio_url
mp3 (default)YesNo
wavYesYes
rawYesYes

Pick encoding and rate for the slices

For wav and raw you can choose encoding from pcm_f32le, pcm_s16le, pcm_mulaw and pcm_alaw, and sample_rate from 8000, 16000, 22050, 24000, 44100 and 48000. The OpenAPI example for avatar muxing uses wav, pcm_s16le and 44100. Gemini's own unary output is a 24 kHz, 16-bit mono WAV (read 2026-10-07), so pcm_s16le at 24000 gives you the same shape if you are swapping engines.

A request that returns clips

Submit as async and poll, because a script of any length can outrun the 30-second sync wait.

curl -X POST https://api.sume.com/v1/tts-1.0/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: slices-001" \
  -d '{
    "transcript": "First line. Second line. Third line.",
    "avatar_handle": "speaker",
    "language": "en",
    "output_format": {"container": "wav", "encoding": "pcm_s16le", "sample_rate": 44100},
    "timestamps": {"words": true},
    "segmentation": {"mode": "sentence"},
    "mode": "async"
  }'

Common mistakes

Wav files are larger than mp3, so if the end product is a web voiceover, request wav only for the cut pass and keep the final export as mp3.

  • Sending segmentation without timestamps.words: true. The schema says segmentation requires it.
  • Expecting slices from the default mp3 and finding only timings.
  • Re-submitting when the sync wait times out. Poll status_url instead; a repeat call with a new key is a second paid job.
  • Cutting the mp3 yourself at the returned times. It works for rough previews, but the slices Sume emits are the sample-exact ones.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume