How do I make a safety briefing audio in three languages with TTS?

One 1,800-character briefing in English, Spanish and Polish is three TTS jobs: 9 cents each, 27 cents in all on Sume. Set language on every job and review it.

4 min readSume
All posts

Make a multilingual safety briefing by translating the script once per language, then sending one TTS job per language with language set to that language's code and a voice that belongs to it. A 1,800-character briefing costs $0.09 per language on Sume ($0.0475 per 1,000 characters, rounded up to the cent), so English, Spanish and Polish together are $0.27. The audio is cheap; the review of the translations is the real work.

This is a worked example for a shop-floor briefing played from a tablet at the start of a shift. Nothing here is legal advice, and a safety instruction that people act on needs a qualified human to approve every language before it is played.

One job per language, never one job with mixed text

The language field tells the voice what language the transcript is in. The API contract asks you to set it for every non-English transcript; if it is omitted the request defaults to English, and Sume only infers Korean or Japanese from a Hangul-only or kana-only transcript. Sending a Spanish script with no language, or with en, is the classic cause of an English-accented Spanish read.

Keep one language per request, because the field names a single language. Split the briefing by language and join the files afterwards if you need a single track; a Timeline audio concat does that for $0.01 without re-synthesis.

Pick a voice that matches, and trust the 409

Every saved voice has a primary language. When your request names a language that the voice does not carry, Sume returns HTTP 409 tts_voice_language_mismatch before it creates a job or charges anything, with the voice's language and the requested one in the body. Retrying the same request with confirm_language_mismatch: true overrides it, but only do that on purpose, for example to test a bilingual voice.

Regional tags compare by primary language, so es-MX and es count as the same language for this check. The check cannot tell you whether the accent suits your workers, so audition one sentence per voice and have a native speaker listen.

What the vendors list, for context

MAI-Voice-2.1's model page says it covers 23 languages, including Spanish, Korean, Turkish, Thai and Vietnamese. Cartesia's languages page lists 50 languages across its platform, including Polish, but does not break the list down by Sonic specifically. ElevenLabs lists 32 languages for Flash, 29 for v2 Multilingual and 70+ for v3. Sume's docs do not publish a per-language list for TTS 1.0, so test each language you need with a real sentence before you plan a rollout, and read the 409 as your coverage check for a given voice.

Language coverage the vendors state on their own pages, read 2026-10-07. Sume publishes no per-language list for TTS 1.0.
ServiceStated coverageSource
MAI-Voice-2.1 and Flash23 languagesmicrosoft.ai model page
Cartesia (platform)50 languages, not split by productcartesia.ai/languages
ElevenLabs Flash / Turbo32 languageselevenlabs.io/pricing/api
ElevenLabs v2 Multilingual29 languageselevenlabs.io/pricing/api
ElevenLabs v370+ languageselevenlabs.io/pricing/api

Submit and price the three jobs

Each language is its own job with its own idempotency key. The script below sends the Polish one; change the transcript, the language code and the key for the others. Use wav if the three files will be joined or if a player needs sample-exact length.

curl -X POST https://api.sume.com/v1/tts-1.0/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: safety-brief-pl-001" \
  -d @- <<JSON
{
  "transcript": "Przed uruchomieniem maszyny sprawdz oslony i wylacznik awaryjny.",
  "voice": { "id": "$VOICE_ID_PL" },
  "language": "pl",
  "generation_config": { "speed": 0.9 },
  "mode": "async"
}
JSON

Cost and review plan

Three languages at 1,800 characters each are $0.27. A retake of one language costs $0.09 again, so a reviewer who catches a mistranslation in the Spanish script only reruns the Spanish job. Slow the voice a little: generation_config.speed accepts 0.6 to 1.5, and 0.9 gives safety text room to land without sounding drawn out.

  • Have a native speaker approve each translated script before generating audio, not after.
  • Write numbers, units and equipment names in the form that is spoken, and check them in the audio.
  • Keep the approved text with the job id so you can show what was played.
  • Re-run all three languages whenever the English source changes.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume