Sume TTS sentence slices need wav or raw output, not mp3
Sume TTS returns per-sentence audio_url slices only for wav or raw output. With mp3 you get timings but no slices. The request that gets clips.

If you ask Sume TTS for sentence segmentation and leave the output at the default, you get timings but no clips. Per-segment audio_url slices exist only when output_format.container is wav or raw. The default is mp3 at 44,100 Hz and 128 kbps, so set the container explicitly when you want one file per sentence.
The rule in the schema
The OpenAPI description for segmentation says it requires timestamps.words: true, returns gapless segments[] (each end equals the next start), and applies a 70 ms post-word boundary by default. emit_audio defaults to true, but "when true and output_format.container is wav/raw, each segment includes a sample-exact audio_url. For mp3, timings are returned without segment audio_url."
That is a deliberate split: lossy mp3 frames do not cut sample-exactly, so Sume only slices PCM.
| Container | Segment timings | Per-segment audio_url |
|---|---|---|
| mp3 (default) | Yes | No |
| wav | Yes | Yes |
| raw | Yes | Yes |
Pick encoding and rate for the slices
For wav and raw you can choose encoding from pcm_f32le, pcm_s16le, pcm_mulaw and pcm_alaw, and sample_rate from 8000, 16000, 22050, 24000, 44100 and 48000. The OpenAPI example for avatar muxing uses wav, pcm_s16le and 44100. Gemini's own unary output is a 24 kHz, 16-bit mono WAV (read 2026-10-07), so pcm_s16le at 24000 gives you the same shape if you are swapping engines.
A request that returns clips
Submit as async and poll, because a script of any length can outrun the 30-second sync wait.
curl -X POST https://api.sume.com/v1/tts-1.0/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: slices-001" \
-d '{
"transcript": "First line. Second line. Third line.",
"avatar_handle": "speaker",
"language": "en",
"output_format": {"container": "wav", "encoding": "pcm_s16le", "sample_rate": 44100},
"timestamps": {"words": true},
"segmentation": {"mode": "sentence"},
"mode": "async"
}'Common mistakes
Wav files are larger than mp3, so if the end product is a web voiceover, request wav only for the cut pass and keep the final export as mp3.
- Sending
segmentationwithouttimestamps.words: true. The schema says segmentation requires it. - Expecting slices from the default mp3 and finding only timings.
- Re-submitting when the sync wait times out. Poll
status_urlinstead; a repeat call with a new key is a second paid job. - Cutting the mp3 yourself at the returned times. It works for rough previews, but the slices Sume emits are the sample-exact ones.
Sources
Related posts
More in Developers
- Sume /v1/usage summary.final is false: a hold is open, not spent
Read GET /v1/usage?job_id= and book cost only when summary.final is true. held_usd_micros and refunded_usd_micros are not spend. Code to poll it.
- Sume video mode: async, sync, subscribe or webhook? Cost is the same
Mode only decides how you learn the outcome: all four create the same job at the same price. A decision table for web apps, workers and batch pipelines.
- Sume webhooks: 10 attempts 30 seconds apart for a video receiver
A Sume job webhook is tried up to 10 times, 30 seconds apart by default, with a 10 s timeout each. What that means for a video receiver, plus a Python verifier.
- Sume webhook.test has no job_id: keep it out of your job table
The dashboard's Send test posts a signed webhook.test with no job_id. Route on event first, dedupe on job_id second, and no phantom job row appears.
Written by Sume