Lobby and elevator announcements with AI voice: 12 messages for $0.09
Twelve recorded building announcements at about 150 characters each cost $0.0855 on Sume TTS. How to batch them, vary languages and reuse the files.

Cost for a full set
A set of 12 building announcements, each about 150 characters, is 1,800 characters. At Sume's $0.0475 per 1,000 characters that is 1.8 x $0.0475 = $0.0855, so about nine cents for the set. You pay once per generation and play the files as often as you like.
Announcements are short, fixed and repeated, which is the best case for text to speech: generate once, review once, store the files.
Plan the messages
List every message with an id before you generate. A lobby usually needs a handful of types, and each type is one job.
| Message | Characters | Cost |
|---|---|---|
| Welcome and visitor check-in | 150 | $0.007125 |
| Elevator out of service | 150 | $0.007125 |
| Fire drill notice | 150 | $0.007125 |
| All 12 messages | 1,800 | $0.0855 |
Generate and store
Submit one POST /v1/tts-1.0/generate per message with a stable Idempotency-Key such as lobby-fire-drill-v1. Poll the job and keep the result URL. If a message text changes, bump the key to v2; the old file stays valid.
For a second language, send the translated text with that language's ISO code in language and a voice that speaks it. Sume's check on language mismatch returns a warning you can confirm, which is a reason to review the first file in each new language before you ship.
Format for your player
The default output is mp3 at 44,100 Hz and 128 kbps, which most signage and speaker systems accept. If your paging system wants a different container or sample rate, the output_format object takes mp3, wav or raw, with sample rates from 8,000 to 48,000 Hz. Test one file on the real hardware before you generate all 12.
- Use
generation_config.speedslightly below 1 for large, echoing rooms. - Keep each message to one idea so staff can swap a single file.
- Version the key and file name together.
Review before you install
Read each script aloud once while you listen to the file. Announcements go wrong in small ways: a room number read as a quantity, an abbreviation spoken letter by letter, a time written as 3:05 that is voiced oddly. Fix these in the text, not in the audio. Write numbers the way you want them spoken, and use a pronunciation dictionary through pronunciation_dict_id for names that recur, such as a tenant or a street.
Keep a table of message id, text, version and job id. When a building rule changes, you will know which files to replace, and you will not have to listen to all twelve again.
Limits worth knowing
A single job can hold up to 20,000 characters and 1,200 seconds of audio, so announcements are nowhere near the ceiling. Do not resubmit a paid request that looks slow; poll the status URL instead, as the jobs guide says.
Sources
Related posts
More in Use cases
- 100 holiday videos in one Sume bulk run: how long with concurrency 16
A bulk queue of 100 items at concurrency 16 drains in 7 waves. Wave arithmetic from the docs' 15 to 30 minute long-form figure, and what else caps the window.
- Dubbed video: burn target-language captions with script_text
After dubbing into Spanish or German, caption the new audio with your translated script as script_text so burned words match your approved translation.
- CAWG identity assertion 1.2: who made it, not whether it is AI
A CAWG identity assertion says which named actor stands behind an asset. It does not say the asset is AI-generated. What Sume stores that you can pair with it.
- CAWG training and data mining 1.1: notAllowed vs constrained
The CAWG training-and-data-mining assertion has four entries, each allowed, notAllowed or constrained. Constrained with no extra info counts as notAllowed.
Written by Sume