Sume TTS default is mp3: sentence slices need wav and emit_audio
Sume TTS defaults to mp3 at 44.1 kHz and 128 kbps. Per-sentence audio slices need wav (pcm_s16le) plus segmentation emit_audio, as in the Python request below.

If you call Sume TTS without an output_format, you get mp3 at 44,100 Hz and 128 kbps. That is fine for a voiceover you play as-is. To get a separate audio file for each sentence you need output_format set to wav with pcm_s16le at 44100, timestamps.words set to true, and segmentation with mode sentence and emit_audio true. With mp3, segment timings come back but not the per-segment audio files.
This is taken from Sume's OpenAPI description of the TTS generate route and the tool guidance published with the hosted MCP server, read on 2026-10-09.
What each setting changes
The request goes to POST /v1/tts-1.0/generate. The transcript can be up to 20,000 characters. You choose a voice with voice.id, or with avatar_id or avatar_handle, in which case Sume resolves the avatar's ready voice. The price is the same whichever format you choose: $0.0475 per 1,000 characters.
| Setting | Default | For sentence slices |
|---|---|---|
| output_format | mp3, 44.1 kHz, 128 kbps | wav, pcm_s16le, 44100 |
| timestamps.words | off | true for word timings |
| segmentation | none | mode sentence, emit_audio true |
| boundary_lead_ms | 70 (inside segmentation) | leave at the default unless a clip cuts too early |
A request that returns slices
The script below sends the request and prints the response. It reads the key and a voice id from environment variables and sends an Idempotency-Key header so a retry does not create a second paid job. The response is a job, so you then poll the job's status and result routes to get the audio files.
import os, requests
body = {
"transcript": "Welcome back. Today we compare three price lists.",
"voice": {"mode": "id", "id": os.environ["SUME_VOICE_ID"]},
"language": "en",
"output_format": {"container": "wav", "sample_rate": 44100,
"encoding": "pcm_s16le"},
"timestamps": {"words": True},
"segmentation": {"mode": "sentence", "emit_audio": True},
}
r = requests.post(
"https://api.sume.com/v1/tts-1.0/generate",
headers={"Authorization": "Bearer " + os.environ["SUME_API_KEY"],
"Idempotency-Key": "tts-slices-001"},
json=body, timeout=30,
)
print(r.status_code)
print(r.json())Common mistakes
Do not send a provider model name or a voice name from another vendor. TTS 1.0 has no engine picker and rejects model and model_id with a 400. A voice id must be a UUID or a voi_ library id; other shapes fail with invalid_voice_id before a job is queued or credits are reserved.
- Keep boundary_lead_ms inside segmentation, not at the top level.
- Do not rebuild the full take by joining the slices; the main audio_url is the continuous take.
- Audio over 1,200 seconds fails with tts_duration_exceeded and no credit is captured.
Where the slices go next
Per-sentence slices are meant to drive other steps. A talking avatar clip can take one sentence of audio at a time, and a timeline can place a shot under each sentence using the word timings. If instead you only need one continuous voiceover under a video, skip segmentation and use the single audio_url from the result, which is the continuous take.
Wav files are larger than mp3 files, so budget for storage if you keep every take. Sume hosts the output on media.sume.com and returns that URL; raw provider URLs are not public outputs.
Cost does not change with the format
The catalog lists TTS at $0.0475 per 1,000 characters with a maximum of 20,000 characters, so a request that you send as mp3 costs the same as one you send as wav. The wav request in this post has 49 characters in its transcript and so costs about a quarter of a cent at that rate, though the catalog also shows a minimum charge of one cent per request.
If you are unsure which format you need, ask for wav with slices once on a short test script. That test costs cents, and it shows exactly what fields come back.
Sources
Related posts
More in Developers
- A 5-minute voice-over: 4.8 MB as 128 kbps MP3, 26.5 MB as WAV
Sume TTS defaults to MP3 at 44,100 Hz and 128 kbps. Five minutes is 4.8 MB; 16-bit mono WAV is 26.46 MB at 44.1 kHz and 28.8 MB at 48 kHz.
- TTS sentence slices: segmentation, wav output and boundary_lead_ms 70
Sume TTS can return gapless sentence segments. Needs timestamps.words and wav or raw for per-sentence audio_url. A 900-character script costs $0.04275.
- Sume TTS sentence segments: boundary_lead_ms 0 vs 500, 12 lines
Ask Sume TTS for words and sentence segmentation and you get gapless wav slices. boundary_lead_ms (default 70) moves each cut 0 to 500 ms after the last word.
- TTS speed 1.2 turns a 30-second read into about 25 seconds
Sume TTS accepts generation_config.speed from 0.6 to 1.5. Speed changes length, not the bill: 450 characters cost $0.021375 at any speed. Test the real length.
Written by Sume