TTS sentence slices: mp3 gives timings only, wav gives audio_urls
Sume TTS segmentation returns sentence timings for any container, but slice audio_urls only with wav or raw. Request shape, the 70 ms rule and when to pick wav.

Sume TTS can cut one voiceover into sentence pieces for you, which saves one job per sentence. There is a catch in the request schema: segmentation.emit_audio defaults to true, but slices are returned as audio only when output_format.container is wav or raw. With mp3, the job returns timings and no per-segment audio_url.
What you must send
timestamps: {words: true}: segmentation requires word timings.segmentation: {mode: "sentence"}: the only mode in v1.output_format: {container: "wav", sample_rate: 44100, encoding: "pcm_s16le"}for audio slices.
What comes back
Segments are gapless: segment[i].end equals segment[i+1].start. A sentence ends 70 ms after its last word by default (boundary_lead_ms, 0 to 500), and the next segment takes the pause. With wav or raw each segment carries a sample-exact audio_url.
Pick the container by the next step
| Next step | Container | Why |
|---|---|---|
| Per-sentence clips for a timeline | wav | Slices get audio_url |
| One narration file for a podcast feed | mp3 | Small file, timings still returned |
| Join takes without gaps later | wav | Sample-domain concat |
| Phone system | wav with pcm_mulaw, 8000 Hz | Telephony format |
A request body
Send it to POST /v1/tts-1.0/generate with your voice selector and an Idempotency-Key.
import json
body = {
"transcript": "First point. Second point. Third point.",
"voice": {"mode": "id", "id": "YOUR_VOICE_ID"},
"timestamps": {"words": True},
"segmentation": {"mode": "sentence", "boundary_lead_ms": 70},
"output_format": {
"container": "wav",
"sample_rate": 44100,
"encoding": "pcm_s16le",
},
}
print(json.dumps(body, indent=2))
Cost
Segmenting does not add a line to the bill; you pay for characters. Three short sentences of 40 characters are 120 characters, 1 cent, the minimum. A 20,000-character script is $0.95 whether or not you segment it.
Related posts
More in Developers
- Pipes and @{} markers in a Sume TTS transcript: stripped, never spoken
Sume strips || cue breaks, @{...} markers and the display side of <display|spoken> before the voice reads; an empty result returns 400 transcript_no_speech.
- TTS with only an API key: list avatars and pick one with voice ready
You do not need a voice id for Sume TTS. List your avatars, pick one whose voice status is ready, and send its handle as avatar_handle. Python example.
- TTS word timings to karaoke captions: map words[] to text
Sume TTS returns words[] with start and end; captions take words as text, start, end. A short Python map skips speech-to-text and burns your exact script.
- Turn STT words into paragraphs: break at pauses over 1.2 seconds
Sume STT returns a flat text string. Use word start and end times to break it into paragraphs at long pauses. Python, no extra API call or cost.
Written by Sume