Sume TTS default is mp3: sentence slices need wav and emit_audio

Sume TTS defaults to mp3 at 44.1 kHz and 128 kbps. Per-sentence audio slices need wav (pcm_s16le) plus segmentation emit_audio, as in the Python request below.

3 min readSume
All posts

If you call Sume TTS without an output_format, you get mp3 at 44,100 Hz and 128 kbps. That is fine for a voiceover you play as-is. To get a separate audio file for each sentence you need output_format set to wav with pcm_s16le at 44100, timestamps.words set to true, and segmentation with mode sentence and emit_audio true. With mp3, segment timings come back but not the per-segment audio files.

This is taken from Sume's OpenAPI description of the TTS generate route and the tool guidance published with the hosted MCP server, read on 2026-10-09.

What each setting changes

The request goes to POST /v1/tts-1.0/generate. The transcript can be up to 20,000 characters. You choose a voice with voice.id, or with avatar_id or avatar_handle, in which case Sume resolves the avatar's ready voice. The price is the same whichever format you choose: $0.0475 per 1,000 characters.

TTS output options as of 2026-10-09 (OpenAPI read 2026-10-09)
SettingDefaultFor sentence slices
output_formatmp3, 44.1 kHz, 128 kbpswav, pcm_s16le, 44100
timestamps.wordsofftrue for word timings
segmentationnonemode sentence, emit_audio true
boundary_lead_ms70 (inside segmentation)leave at the default unless a clip cuts too early

A request that returns slices

The script below sends the request and prints the response. It reads the key and a voice id from environment variables and sends an Idempotency-Key header so a retry does not create a second paid job. The response is a job, so you then poll the job's status and result routes to get the audio files.

import os, requests

body = {
    "transcript": "Welcome back. Today we compare three price lists.",
    "voice": {"mode": "id", "id": os.environ["SUME_VOICE_ID"]},
    "language": "en",
    "output_format": {"container": "wav", "sample_rate": 44100,
                      "encoding": "pcm_s16le"},
    "timestamps": {"words": True},
    "segmentation": {"mode": "sentence", "emit_audio": True},
}
r = requests.post(
    "https://api.sume.com/v1/tts-1.0/generate",
    headers={"Authorization": "Bearer " + os.environ["SUME_API_KEY"],
             "Idempotency-Key": "tts-slices-001"},
    json=body, timeout=30,
)
print(r.status_code)
print(r.json())

Common mistakes

Do not send a provider model name or a voice name from another vendor. TTS 1.0 has no engine picker and rejects model and model_id with a 400. A voice id must be a UUID or a voi_ library id; other shapes fail with invalid_voice_id before a job is queued or credits are reserved.

  • Keep boundary_lead_ms inside segmentation, not at the top level.
  • Do not rebuild the full take by joining the slices; the main audio_url is the continuous take.
  • Audio over 1,200 seconds fails with tts_duration_exceeded and no credit is captured.

Where the slices go next

Per-sentence slices are meant to drive other steps. A talking avatar clip can take one sentence of audio at a time, and a timeline can place a shot under each sentence using the word timings. If instead you only need one continuous voiceover under a video, skip segmentation and use the single audio_url from the result, which is the continuous take.

Wav files are larger than mp3 files, so budget for storage if you keep every take. Sume hosts the output on media.sume.com and returns that URL; raw provider URLs are not public outputs.

Cost does not change with the format

The catalog lists TTS at $0.0475 per 1,000 characters with a maximum of 20,000 characters, so a request that you send as mp3 costs the same as one you send as wav. The wav request in this post has 49 characters in its transcript and so costs about a quarter of a cent at that rate, though the catalog also shows a minimum charge of one cent per request.

If you are unsure which format you need, ask for wav with slices once on a short test script. That test costs cents, and it shows exactly what fields come back.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume