How do I make school morning announcements audio with an AI voice?
A 1,500-character daily announcement is 8 cents on Sume TTS, about $14.40 for 180 school days. How to template the script, add Spanish and keep it short.

To make school morning announcements with an AI voice, keep a fixed template, fill in the day's items as text, and send the finished script to a TTS endpoint as one job. A 1,500-character script is about two minutes of speech and costs $0.08 on Sume TTS 1.0, because billing rounds each job up to whole cents. Over 180 school days that is $14.40.
The script is the work here, not the audio. Announcements repeat the same shape every morning: greeting, date, lunch, a club reminder, one thing to be proud of, a sign-off. That structure is what lets you automate the audio without it sounding like a robot reading a list.
A template that stays under two minutes
Cartesia's pricing page puts a minute of Sonic speech at 750 to 800 credits, one credit per character, so 1,500 characters lands near two minutes. Students stop listening past that. Cap each slot so the total never passes the budget.
- Greeting and date: about 120 characters.
- Lunch and schedule changes: about 300 characters.
- Clubs and sports, no more than three items: about 450 characters.
- One recognition or reminder: about 400 characters.
- Sign-off: about 80 characters.
Write for the ear, then check the voice
Spell dates the way they should be said ('Tuesday, October thirteenth') rather than relying on a numeric form. Write room numbers and club acronyms with spaces or hyphens if they are read badly, and keep a pronunciation dictionary for staff and place names so the same fix applies every morning. Sume TTS accepts a pronunciation_dict_id on the request.
Set the tone with generation_config.emotion, a free-text field up to 64 characters, and nudge speed between 0.6 and 1.5 if a slower read suits younger students. Both settings change how the voice sounds, not the price, which depends only on characters.
One job a day, run from a schedule
Send the script as an async job and pass a webhook or poll the job, then fetch the result when the job is terminal. Use a date-based Idempotency-Key such as announce-2026-10-13 so a retry cannot generate and bill a second copy. Posting to the school's player is then a download and an upload.
curl -X POST https://api.sume.com/v1/tts-1.0/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: announce-2026-10-13" \
-d '{
"transcript": "Good morning, everyone. It is Tuesday, October thirteenth. Today is pasta day in the cafeteria.",
"voice": { "id": "'"$VOICE_ID"'" },
"language": "en",
"mode": "async"
}'Adding Spanish for families
A second language is a second job, not a surcharge: translate the same template, send it with language set to es, and the cost is one more job of the same length, $0.08 a day or $28.80 across 180 days for both. Choose a voice made for the language; if the voice and language disagree, Sume returns a 409 tts_voice_language_mismatch before charging, and you can confirm it deliberately with confirm_language_mismatch.
Keeping it listenable all year
The risk with a daily read is sameness. Rotate the sign-off line weekly, let a student write one item a week and run it through the same template, and keep a short list of banned phrases such as 'please be advised'. A person should read the script before it is submitted: a mistake in a date is a retake, and a retake is another job at the same rate, so a proofread is cheaper than a regenerate.
If you want a jingle at the start, generate it once as a music track, keep the file, and join it to each morning's voice with a Timeline audio concat, which is a flat $0.01 a job and does not re-synthesize anything.
The price next to the vendors
The table uses the same assumption for each: 1,500 characters a day, 180 days, 270,000 characters a year, with no per-job rounding on the vendor side because their pages list per-character rates only.
- Every option is a few dollars a year for the whole school, so choose on voice quality and languages, not price.
- The Sume figure includes whole-cent rounding per job; with 1,500 characters that is the difference between $0.071 and $0.08 a day.
| Provider and model | Listed rate | Year of announcements |
|---|---|---|
| Sume TTS 1.0 | $0.0475 per 1,000 characters | $14.40 |
| Microsoft MAI-Voice-2.1 Standard | $22 per 1M characters | $5.94 |
| Microsoft MAI-Voice-2.1 Flash | $15 per 1M characters | $4.05 |
| ElevenLabs Flash/Turbo | $0.04 per 1,000 characters | $10.80 |
| OpenAI tts-1 | $15 per 1M characters | $4.05 |
Sources
Related posts
More in Use cases
- How do I make a self-guided walking tour audio with TTS?
A six-stop walking tour at 700 characters a stop is six TTS jobs of 4 cents each, 24 cents on Sume, about 0.9 MB per stop as default mp3.
- How do I add a spoken tagline and sound sting to the end of an ad?
Make a 3-second spoken end tag: one 1-cent TTS job, a $0.125 music sting and a $0.01 concat to attach it to any ad read. About 14.5 cents on Sume.
- Swap the presenter per market: one ad, two Recast jobs on Sume
Recast keeps the motion, camera, cuts and sound of one source ad and replaces the people. How to make one version per market, the limits and the price.
- 10 static ad variants in one Sume image request: n and its cap
The Sume Image API takes n from 1 to 10, but each model has its own cap. Read it from the catalog; use 4:5 on Nano Banana or 1088 x 1360 on GPT for feed ads.
Written by Sume