One TTS job with sentence segments vs one job per sentence

Sume TTS can return gapless sentence timings and per-sentence wav slices from one job. The price per character is the same; the jobs, keys and polls are not.

4 min readSume
All posts

If you need a voiceover as separate sentence files, make one Sume TTS job with timestamps: {words: true} and segmentation: {mode: "sentence"} instead of one job per sentence. The character price is the same, $0.0475 per 1,000, but one job means one idempotency key, one poll and one result. The result returns gapless segments[] and, for wav or raw output, a sample-exact audio_url for each (API reference).

What the single job gives you

  • Gapless segments: segments[i].end equals segments[i+1].start.
  • A cut that follows the last word of each sentence by boundary_lead_ms, 70 ms by default, in a range of 0 to 500. The next segment absorbs the pause.
  • Slices only for wav and raw. With mp3 you get timings without per-segment audio_url.
  • emit_audio defaults to true, so request wav and you get the slices.
{
  "transcript": "Hook line. Offer line. Call to action.",
  "avatar_handle": "your-voice-handle",
  "output_format": {"container": "wav", "encoding": "pcm_s16le", "sample_rate": 44100},
  "timestamps": {"words": true},
  "segmentation": {"mode": "sentence", "boundary_lead_ms": 70}
}

The case for separate jobs

One job per sentence is still right when you want to redo a line without touching the rest. A change in one sentence of a single job means a new job for the whole script. At 1,000 characters that is $0.0475, so the extra cost of a full redo is small, and the new take may differ in delivery from the earlier one. Per-line jobs keep the untouched sentences exactly as they were.

Compare the two

Sentence files from one job versus many, at the same $0.0475 per 1,000 characters (read 2026-10-04)
QuestionOne job with segmentationOne job per sentence
Idempotency keys1one per sentence
Polls and results1one per sentence
Pacing across sentencesSentences rendered togetherEach sentence rendered alone
Redo one sentenceNew job for the full scriptOne small job
Max per request20,000 characters20,000 characters each

The trend behind it

Streaming speech is the headline: Microsoft quotes MAI-Voice-2.1-Flash at 150 ms end to end for 45 seconds of audio (Microsoft AI, read 2026-10-04). If you want sentence-level control from a job API, segmentation is the equivalent. Poll the job as the jobs docs describe and read segments[] when the result is ready.

Rule of thumb: use one job with segments for a script that is read in order, and per-sentence jobs for lines that are edited and swapped on their own.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume