How do I make a safety briefing audio in three languages with TTS?
One 1,800-character briefing in English, Spanish and Polish is three TTS jobs: 9 cents each, 27 cents in all on Sume. Set language on every job and review it.

Make a multilingual safety briefing by translating the script once per language, then sending one TTS job per language with language set to that language's code and a voice that belongs to it. A 1,800-character briefing costs $0.09 per language on Sume ($0.0475 per 1,000 characters, rounded up to the cent), so English, Spanish and Polish together are $0.27. The audio is cheap; the review of the translations is the real work.
This is a worked example for a shop-floor briefing played from a tablet at the start of a shift. Nothing here is legal advice, and a safety instruction that people act on needs a qualified human to approve every language before it is played.
One job per language, never one job with mixed text
The language field tells the voice what language the transcript is in. The API contract asks you to set it for every non-English transcript; if it is omitted the request defaults to English, and Sume only infers Korean or Japanese from a Hangul-only or kana-only transcript. Sending a Spanish script with no language, or with en, is the classic cause of an English-accented Spanish read.
Keep one language per request, because the field names a single language. Split the briefing by language and join the files afterwards if you need a single track; a Timeline audio concat does that for $0.01 without re-synthesis.
Pick a voice that matches, and trust the 409
Every saved voice has a primary language. When your request names a language that the voice does not carry, Sume returns HTTP 409 tts_voice_language_mismatch before it creates a job or charges anything, with the voice's language and the requested one in the body. Retrying the same request with confirm_language_mismatch: true overrides it, but only do that on purpose, for example to test a bilingual voice.
Regional tags compare by primary language, so es-MX and es count as the same language for this check. The check cannot tell you whether the accent suits your workers, so audition one sentence per voice and have a native speaker listen.
What the vendors list, for context
MAI-Voice-2.1's model page says it covers 23 languages, including Spanish, Korean, Turkish, Thai and Vietnamese. Cartesia's languages page lists 50 languages across its platform, including Polish, but does not break the list down by Sonic specifically. ElevenLabs lists 32 languages for Flash, 29 for v2 Multilingual and 70+ for v3. Sume's docs do not publish a per-language list for TTS 1.0, so test each language you need with a real sentence before you plan a rollout, and read the 409 as your coverage check for a given voice.
| Service | Stated coverage | Source |
|---|---|---|
| MAI-Voice-2.1 and Flash | 23 languages | microsoft.ai model page |
| Cartesia (platform) | 50 languages, not split by product | cartesia.ai/languages |
| ElevenLabs Flash / Turbo | 32 languages | elevenlabs.io/pricing/api |
| ElevenLabs v2 Multilingual | 29 languages | elevenlabs.io/pricing/api |
| ElevenLabs v3 | 70+ languages | elevenlabs.io/pricing/api |
Submit and price the three jobs
Each language is its own job with its own idempotency key. The script below sends the Polish one; change the transcript, the language code and the key for the others. Use wav if the three files will be joined or if a player needs sample-exact length.
curl -X POST https://api.sume.com/v1/tts-1.0/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: safety-brief-pl-001" \
-d @- <<JSON
{
"transcript": "Przed uruchomieniem maszyny sprawdz oslony i wylacznik awaryjny.",
"voice": { "id": "$VOICE_ID_PL" },
"language": "pl",
"generation_config": { "speed": 0.9 },
"mode": "async"
}
JSONCost and review plan
Three languages at 1,800 characters each are $0.27. A retake of one language costs $0.09 again, so a reviewer who catches a mistranslation in the Spanish script only reruns the Spanish job. Slow the voice a little: generation_config.speed accepts 0.6 to 1.5, and 0.9 gives safety text room to land without sounding drawn out.
- Have a native speaker approve each translated script before generating audio, not after.
- Write numbers, units and equipment names in the form that is spoken, and check them in the audio.
- Keep the approved text with the job id so you can show what was played.
- Re-run all three languages whenever the English source changes.
Sources
Related posts
More in Use cases
- How do I narrate a family history video with AI voice and photos?
Narrate a family history from old photos: 1,800 characters of script is $0.09 on Sume TTS, plus a $0.125 music bed and a $0.30 Timeline render.
- What does a first-time home buyer video series cost with an avatar?
Six 30-second avatar tips cost $33.12 at standard, $44.10 at plus or $99.00 at max on Sume, plus $0.95 once for the avatar. Rates and script limits.
- Four seasons from one venue photo: four Ideogram 4.5 edits for $0.30
Make spring, summer, autumn and winter versions of one venue photo with four ideogram/ideogram-v4.5 edits at $0.075 each on Sume, $7.50 for 25 venues.
- How do I turn my free course lessons into text for a sales page?
Detach each lesson's audio ($0.01), then transcribe it at $0.01 a minute: $0.09 for an 8-minute lesson, $0.90 for ten, on Sume.
Written by Sume