Dictation audio: 12 sentences in one TTS job, 6 cents, wav slices

Make 12 dictation sentences as separate wav files from one Sume TTS job with sentence segmentation. 1,080 characters cost 6 cents, against 12 cents as 12 jobs.

5 min readSume
All posts

Short answer

Send all 12 sentences as one transcript to Sume TTS with timestamps.words: true and segmentation: { mode: "sentence" }, and ask for a wav container. Each sentence then comes back as its own sample-exact audio_url. At 90 characters per sentence the job is 1,080 characters, which is 5.13 cents at $47.50 per million and rounds up to 6 cents. Twelve separate jobs would each round up to 1 cent, so 12 cents.

The segmentation fields are from the Sume OpenAPI reference: segmentation needs word timestamps, returns gapless segments[], and emit_audio slices need a wav or raw container; with mp3 you get timings but no per-segment audio_url.

A teacher wants each sentence as a separate file for a listening dictation, with a pause that students fill in by writing. One request is simpler to track than twelve, and the cost drops because the per-job rounding applies once.

Two ways to make 12 sentences of 90 characters (Sume TTS public rate, read 2026-10-07)
ApproachCharacters billedUnrounded costBilled
One job, sentence segments1,080$0.0513$0.06
Twelve jobs12 x 90 = 1,080$0.0043 each$0.12

Check the arithmetic

Cost per job is characters times $47.50 per million, rounded up to a whole cent. Spaces and punctuation count. One job must stay within 20,000 characters and 1,200 seconds, which 12 short sentences are nowhere near.

The request

Keep the sentences in one string, separated by a space, each ending with a full stop so the sentence boundary is clear. The job result lists segments[] in order, so segment 0 is sentence 1.

{
  "transcript": "The train leaves at nine. Please bring your passport. ...",
  "voice": { "mode": "id", "id": "<voice id>" },
  "output_format": { "container": "wav", "encoding": "pcm_s16le", "sample_rate": 44100 },
  "timestamps": { "words": true },
  "segmentation": { "mode": "sentence", "emit_audio": true }
}

Pauses between sentences

Segments are gapless: each one ends where the next begins, and by default the cut falls 70 ms after the last word, so the next segment absorbs the pause. The boundary_lead_ms field accepts 0 to 500 if you want a different cut point. For a longer writing gap, add the pause in your player or app between files. A timeline audio concat job can join the slices into one file if you prefer that.

When not to use one job

Use mp3 if you only need the timings. The segments then have no individual files, so for a dictation that plays sentence by sentence, stay on wav.

Repeats

If a student must hear a sentence twice, play the same file twice rather than generating it again, which would bill again.

Scaling to a week of exercises

The one-job approach keeps its advantage as a set grows, up to the per-job limits of 20,000 characters and 1,200 seconds. A week of five dictations, each 12 sentences of 90 characters, is five jobs of 1,080 characters, which is 30 cents, against 60 cents as single-sentence jobs.

Keep a small table of your own with exercise, job id and character count. The job result lists segments[] in sentence order, so a missing or extra sentence shows up as a wrong count before a student ever hears the file.

If an exercise has a sentence you need to change, regenerate the exercise as a whole, since the segments are cut from one rendering and neighbors can shift.

  • Count characters with JavaScript string length rules, as Sume does.
  • Stay on wav when you need per-sentence files.
  • Regenerate the whole exercise for one edited sentence.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume