How do I make school morning announcements audio with an AI voice?

A 1,500-character daily announcement is 8 cents on Sume TTS, about $14.40 for 180 school days. How to template the script, add Spanish and keep it short.

4 min readSume
All posts

To make school morning announcements with an AI voice, keep a fixed template, fill in the day's items as text, and send the finished script to a TTS endpoint as one job. A 1,500-character script is about two minutes of speech and costs $0.08 on Sume TTS 1.0, because billing rounds each job up to whole cents. Over 180 school days that is $14.40.

The script is the work here, not the audio. Announcements repeat the same shape every morning: greeting, date, lunch, a club reminder, one thing to be proud of, a sign-off. That structure is what lets you automate the audio without it sounding like a robot reading a list.

A template that stays under two minutes

Cartesia's pricing page puts a minute of Sonic speech at 750 to 800 credits, one credit per character, so 1,500 characters lands near two minutes. Students stop listening past that. Cap each slot so the total never passes the budget.

  • Greeting and date: about 120 characters.
  • Lunch and schedule changes: about 300 characters.
  • Clubs and sports, no more than three items: about 450 characters.
  • One recognition or reminder: about 400 characters.
  • Sign-off: about 80 characters.

Write for the ear, then check the voice

Spell dates the way they should be said ('Tuesday, October thirteenth') rather than relying on a numeric form. Write room numbers and club acronyms with spaces or hyphens if they are read badly, and keep a pronunciation dictionary for staff and place names so the same fix applies every morning. Sume TTS accepts a pronunciation_dict_id on the request.

Set the tone with generation_config.emotion, a free-text field up to 64 characters, and nudge speed between 0.6 and 1.5 if a slower read suits younger students. Both settings change how the voice sounds, not the price, which depends only on characters.

One job a day, run from a schedule

Send the script as an async job and pass a webhook or poll the job, then fetch the result when the job is terminal. Use a date-based Idempotency-Key such as announce-2026-10-13 so a retry cannot generate and bill a second copy. Posting to the school's player is then a download and an upload.

curl -X POST https://api.sume.com/v1/tts-1.0/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: announce-2026-10-13" \
  -d '{
    "transcript": "Good morning, everyone. It is Tuesday, October thirteenth. Today is pasta day in the cafeteria.",
    "voice": { "id": "'"$VOICE_ID"'" },
    "language": "en",
    "mode": "async"
  }'

Adding Spanish for families

A second language is a second job, not a surcharge: translate the same template, send it with language set to es, and the cost is one more job of the same length, $0.08 a day or $28.80 across 180 days for both. Choose a voice made for the language; if the voice and language disagree, Sume returns a 409 tts_voice_language_mismatch before charging, and you can confirm it deliberately with confirm_language_mismatch.

Keeping it listenable all year

The risk with a daily read is sameness. Rotate the sign-off line weekly, let a student write one item a week and run it through the same template, and keep a short list of banned phrases such as 'please be advised'. A person should read the script before it is submitted: a mistake in a date is a retake, and a retake is another job at the same rate, so a proofread is cheaper than a regenerate.

If you want a jingle at the start, generate it once as a music track, keep the file, and join it to each morning's voice with a Timeline audio concat, which is a flat $0.01 a job and does not re-synthesize anything.

The price next to the vendors

The table uses the same assumption for each: 1,500 characters a day, 180 days, 270,000 characters a year, with no per-job rounding on the vendor side because their pages list per-character rates only.

  • Every option is a few dollars a year for the whole school, so choose on voice quality and languages, not price.
  • The Sume figure includes whole-cent rounding per job; with 1,500 characters that is the difference between $0.071 and $0.08 a day.
Annual cost for 270,000 characters of announcements, vendor list rates read 2026-10-07 and Sume TTS 1.0 catalog rate.
Provider and modelListed rateYear of announcements
Sume TTS 1.0$0.0475 per 1,000 characters$14.40
Microsoft MAI-Voice-2.1 Standard$22 per 1M characters$5.94
Microsoft MAI-Voice-2.1 Flash$15 per 1M characters$4.05
ElevenLabs Flash/Turbo$0.04 per 1,000 characters$10.80
OpenAI tts-1$15 per 1M characters$4.05

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume