One TTS job with sentence segments vs one job per sentence
Sume TTS can return gapless sentence timings and per-sentence wav slices from one job. The price per character is the same; the jobs, keys and polls are not.

If you need a voiceover as separate sentence files, make one Sume TTS job with timestamps: {words: true} and segmentation: {mode: "sentence"} instead of one job per sentence. The character price is the same, $0.0475 per 1,000, but one job means one idempotency key, one poll and one result. The result returns gapless segments[] and, for wav or raw output, a sample-exact audio_url for each (API reference).
What the single job gives you
- Gapless segments:
segments[i].endequalssegments[i+1].start. - A cut that follows the last word of each sentence by
boundary_lead_ms, 70 ms by default, in a range of 0 to 500. The next segment absorbs the pause. - Slices only for
wavandraw. With mp3 you get timings without per-segmentaudio_url. emit_audiodefaults to true, so request wav and you get the slices.
{
"transcript": "Hook line. Offer line. Call to action.",
"avatar_handle": "your-voice-handle",
"output_format": {"container": "wav", "encoding": "pcm_s16le", "sample_rate": 44100},
"timestamps": {"words": true},
"segmentation": {"mode": "sentence", "boundary_lead_ms": 70}
}The case for separate jobs
One job per sentence is still right when you want to redo a line without touching the rest. A change in one sentence of a single job means a new job for the whole script. At 1,000 characters that is $0.0475, so the extra cost of a full redo is small, and the new take may differ in delivery from the earlier one. Per-line jobs keep the untouched sentences exactly as they were.
Compare the two
| Question | One job with segmentation | One job per sentence |
|---|---|---|
| Idempotency keys | 1 | one per sentence |
| Polls and results | 1 | one per sentence |
| Pacing across sentences | Sentences rendered together | Each sentence rendered alone |
| Redo one sentence | New job for the full script | One small job |
| Max per request | 20,000 characters | 20,000 characters each |
The trend behind it
Streaming speech is the headline: Microsoft quotes MAI-Voice-2.1-Flash at 150 ms end to end for 45 seconds of audio (Microsoft AI, read 2026-10-04). If you want sentence-level control from a job API, segmentation is the equivalent. Poll the job as the jobs docs describe and read segments[] when the result is ready.
Rule of thumb: use one job with segments for a script that is read in order, and per-sentence jobs for lines that are edited and swapped on their own.
Sources
Related posts
More in Developers
- OpenAI Agents API hosted sandbox: which Sume hosts to allow
OpenAI's Agents API is in public beta with hosted or connected sandboxes. Which Sume hosts to allow, how to pass the MCP URL, and why the key stays in a secret.
- OpenAI Agents API sandbox: keep the Sume API key out of it
The OpenAI Agents API beta adds a sandbox and hosted-browser computer use. Where a Sume API key can live when an agent runs there, and what to hand it instead.
- OpenAI named no Videos API replacement: build it swappable
OpenAI's deprecation page names no successor to the Sora Videos API. Put video behind one interface, discover models from the catalog, and keep ids out of code.
- Mask file checklist for GPT Image 2.5: alpha, same size, under 50 MB
OpenAI says an edit mask must match the image in format and size, stay under 50 MB, and carry an alpha channel. Check your file before sending mask_url to Sume.
Written by Sume