How do I make a bilingual English and Spanish audio announcement?

Make one bilingual announcement file: two TTS jobs, one per language, joined by a $0.01 Timeline audio concat with no re-synthesis. About 10 cents in total.

4 min readSume
All posts

To make a bilingual announcement as one audio file, generate each language as its own TTS job with the right voice and language, then join the two files with a Timeline audio concat. A 380-character English read costs $0.02, a 420-character Spanish read costs $0.02, and the concat is a flat $0.01, so $0.05 in total on Sume.

Two jobs beat one mixed-language script. A voice built for English reading Spanish words usually sounds wrong in both; separate jobs let you pick a voice that suits each language and keep the timing of each part under your control.

Why not one script with both languages

Sume infers only Korean and Japanese from the text, and only when it is Hangul-only or kana-only. For anything else, including English and Spanish, you must send language yourself, so a single request cannot switch language mid-way. Each job carries one language and one voice.

There is also a guard. If you send a voice that is tagged for one language with a transcript marked as another, Sume returns a 409 tts_voice_language_mismatch before any charge. Regional tags are compared by primary language, so es-MX and es do not clash. Retry with confirm_language_mismatch: true only when you mean it.

The two jobs

Send both with distinct idempotency keys and output as wav, so the concat has clean input. Use the same speed in both languages unless you have a reason not to; Count the characters in each version, since the cost follows the text you send.

for LANG in en es; do
  curl -s -X POST https://api.sume.com/v1/tts-1.0/generate \
    -H "Authorization: Bearer $SUME_API_KEY" \
    -H "Content-Type: application/json" \
    -H "Idempotency-Key: pa-notice-$LANG-1" \
    -d "{\"transcript\": \"$(cat notice.$LANG.txt)\", \"voice\": {\"id\": \"$(printenv VOICE_$LANG)\"}, \"language\": \"$LANG\", \"output_format\": {\"container\": \"wav\", \"encoding\": \"pcm_s16le\", \"sample_rate\": 44100}}"
done

Join them with a concat

Timeline 1.0 audio takes 1 to 20 parts and joins them gapless without re-synthesis, at a flat $0.01 a job. Put a short chime or a beat of silence between the two languages if you want a clear break; that is just another part. Output is wav or mp3, up to 1,800 seconds, and mp3 adds a little encoder padding, so pick wav if you will concat again.

curl -X POST https://api.sume.com/v1/timeline-1.0/audio \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: pa-notice-joined-1" \
  -d '{
    "operation": "concat",
    "parts": [{ "url": "'"$EN_URL"'" }, { "url": "'"$ES_URL"'" }],
    "output": { "format": "mp3" }
  }'

Writing both halves

Write the two versions together, not one after the other. Start from the facts: what happens, where, when, and what a listener should do. Then say each fact once in each language, in the same order, so a listener can follow either half. Keep numbers, times and places as spoken words, and spell out street names the way they are said.

Check pronunciation of names that appear in both halves. A place name read in an English accent inside the Spanish half is jarring. Sume TTS accepts a pronunciation_dict_id per request, so keep one dictionary per language and attach it to the matching job.

Finally, set the lengths. If one half runs much longer, listeners in the other language may think the message has finished. A second or two of difference is fine; thirty seconds is not.

When to use separate files instead

Sometimes one file is the wrong answer. If the announcement plays on two channels, or visitors choose their language at a kiosk, keep the two files separate and skip the concat. Nothing is lost: the jobs are already the same price. The concat earns its $0.01 only when a single stream has to carry both languages, and it keeps the segment offsets, so you can still find where the second half starts.

Order, length and price

Say the language you expect most people to need first, and keep both versions equal in content, not just in topic. Public announcements are often safety-related, so have a fluent speaker check the translation and the final audio before it plays.

The cost is small, so the real saving is time. A change to one notice is a new job in one language plus a concat, not a re-record of the whole announcement.

Cost of an 800-character bilingual announcement, Sume TTS 1.0 catalog and Timeline audio rates as of 2026-10-07; vendor per-character list rates read 2026-10-07 for comparison.
ProviderListed rate800 characters
Sume TTS 1.0, two jobs plus concat$0.0475 per 1,000 characters, $0.01 per concat$0.05
ElevenLabs Flash/Turbo$0.04 per 1,000 characters$0.032
ElevenLabs v3$0.08 per 1,000 characters$0.064

Sources

Related posts

More in Developers

All Developers posts

Written by Sume