Game NPC barks: 150 short lines in one TTS job (43 cents) vs 150 jobs
Short lines cost 1 cent each as separate Sume TTS jobs. One job with sentence slices returns the same 150 lines as separate WAV files for 43 cents.

Sume TTS rounds each job up to a whole cent with a 1-cent minimum, so 150 barks of 60 characters cost 150 cents as 150 jobs. Put them in one script and request sentence segmentation, and the same lines come back as 150 sample-exact WAV slices for about 43 cents. The saving comes entirely from not paying the minimum 150 times.
Why do tiny jobs cost more per character?
The rate is $0.0475 per 1,000 characters. A 60-character line is 0.285 cents, which rounds up to 1 cent alone. Anything up to 210 characters rounds to 1 cent as its own job, so short lines waste most of what they pay for.
| Approach | Jobs | Characters billed | Charge |
|---|---|---|---|
| One job per line | 150 | 9,000 total | 150 cents |
| One job, all lines | 1 | 9,000 | 43 cents ($0.4275 rounded up) |
How do I get separate files back?
Send the lines as one transcript, one sentence per line with terminal punctuation. Ask for timestamps: { words: true }, segmentation: { mode: "sentence" } and a wav container. Per the OpenAPI schema, segmentation requires word timestamps, returns gapless segments[], and with a wav or raw container each segment includes a sample-exact audio_url. With mp3 you get timings but no per-segment files.
Segments follow the sentence boundaries of the script, in order, so segment N is line N as long as each line ends in its own full stop, question mark or exclamation mark. Abbreviations can split a sentence early, as covered in the abbreviation post. Compare the segment count with your line count before importing files into the game.
What are the trade-offs?
Everything in one job shares one voice, language and speed. Barks for a different character need a different job. Also keep the whole script within 20,000 characters, which is about 330 lines of this size. One bad line means regenerating the batch or doing a 1-cent single-line job for the fix. Per-job cost applies again at that point, and that is fine for a one-off.
The default pause rule gives each segment the 70 ms after its last word, and boundary_lead_ms (0 to 500) changes it. Trim any tail you do not want inside your game engine.
Sources
Related posts
More in Developers
- gemini-3.1-flash-image deprecated: Nano Banana 2.1 on Sume
Google deprecated gemini-3.1-flash-image on Oct 6 when Nano Banana 2.1 went GA. Which ids Sume accepts, what it bills per image, and what Google lists.
- Gemini 3.7 Flash now routes to 3.8: where a media client pins ids
Google auto-routes gemini-3.7-flash to gemini-3.8-flash since Oct 8. Where Sume lets you pin a model, where it routes for you, and what the job records.
- Gemini Deep Research agent shuts down Oct 23: an async alternative
Google marked deep-research-pro-preview-12-2025 for shutdown on Oct 23, 2026. What an async Sume Agent Completion run does, and what it does not do like it.
- Gemini Omni Flash 1.1 prompt length on Sume: a 20,000-character check
The Sume catalog constraint for Gemini Omni Flash 1.1 is a prompt of at most 20,000 characters, 3 to 10 seconds. A short Python preflight check before you pay.
Written by Sume