How do I make a bilingual English and Spanish audio announcement?
Make one bilingual announcement file: two TTS jobs, one per language, joined by a $0.01 Timeline audio concat with no re-synthesis. About 10 cents in total.

To make a bilingual announcement as one audio file, generate each language as its own TTS job with the right voice and language, then join the two files with a Timeline audio concat. A 380-character English read costs $0.02, a 420-character Spanish read costs $0.02, and the concat is a flat $0.01, so $0.05 in total on Sume.
Two jobs beat one mixed-language script. A voice built for English reading Spanish words usually sounds wrong in both; separate jobs let you pick a voice that suits each language and keep the timing of each part under your control.
Why not one script with both languages
Sume infers only Korean and Japanese from the text, and only when it is Hangul-only or kana-only. For anything else, including English and Spanish, you must send language yourself, so a single request cannot switch language mid-way. Each job carries one language and one voice.
There is also a guard. If you send a voice that is tagged for one language with a transcript marked as another, Sume returns a 409 tts_voice_language_mismatch before any charge. Regional tags are compared by primary language, so es-MX and es do not clash. Retry with confirm_language_mismatch: true only when you mean it.
The two jobs
Send both with distinct idempotency keys and output as wav, so the concat has clean input. Use the same speed in both languages unless you have a reason not to; Count the characters in each version, since the cost follows the text you send.
for LANG in en es; do
curl -s -X POST https://api.sume.com/v1/tts-1.0/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: pa-notice-$LANG-1" \
-d "{\"transcript\": \"$(cat notice.$LANG.txt)\", \"voice\": {\"id\": \"$(printenv VOICE_$LANG)\"}, \"language\": \"$LANG\", \"output_format\": {\"container\": \"wav\", \"encoding\": \"pcm_s16le\", \"sample_rate\": 44100}}"
doneJoin them with a concat
Timeline 1.0 audio takes 1 to 20 parts and joins them gapless without re-synthesis, at a flat $0.01 a job. Put a short chime or a beat of silence between the two languages if you want a clear break; that is just another part. Output is wav or mp3, up to 1,800 seconds, and mp3 adds a little encoder padding, so pick wav if you will concat again.
curl -X POST https://api.sume.com/v1/timeline-1.0/audio \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: pa-notice-joined-1" \
-d '{
"operation": "concat",
"parts": [{ "url": "'"$EN_URL"'" }, { "url": "'"$ES_URL"'" }],
"output": { "format": "mp3" }
}'Writing both halves
Write the two versions together, not one after the other. Start from the facts: what happens, where, when, and what a listener should do. Then say each fact once in each language, in the same order, so a listener can follow either half. Keep numbers, times and places as spoken words, and spell out street names the way they are said.
Check pronunciation of names that appear in both halves. A place name read in an English accent inside the Spanish half is jarring. Sume TTS accepts a pronunciation_dict_id per request, so keep one dictionary per language and attach it to the matching job.
Finally, set the lengths. If one half runs much longer, listeners in the other language may think the message has finished. A second or two of difference is fine; thirty seconds is not.
When to use separate files instead
Sometimes one file is the wrong answer. If the announcement plays on two channels, or visitors choose their language at a kiosk, keep the two files separate and skip the concat. Nothing is lost: the jobs are already the same price. The concat earns its $0.01 only when a single stream has to carry both languages, and it keeps the segment offsets, so you can still find where the second half starts.
Order, length and price
Say the language you expect most people to need first, and keep both versions equal in content, not just in topic. Public announcements are often safety-related, so have a fluent speaker check the translation and the final audio before it plays.
The cost is small, so the real saving is time. A change to one notice is a new job in one language plus a concat, not a re-record of the whole announcement.
| Provider | Listed rate | 800 characters |
|---|---|---|
| Sume TTS 1.0, two jobs plus concat | $0.0475 per 1,000 characters, $0.01 per concat | $0.05 |
| ElevenLabs Flash/Turbo | $0.04 per 1,000 characters | $0.032 |
| ElevenLabs v3 | $0.08 per 1,000 characters | $0.064 |
Sources
Related posts
More in Developers
- callback_url or webhook_url: which field each Sume video route takes
POST /v1/videos takes callback_url; motion control, lip-sync and image routes take mode plus webhook_url. The field names and what they share.
- Cap Sume spend from an agent loop: dry_run, max_spend_usd and run caps
Four optional guards cap what an automated Sume caller can spend: dry_run, max_spend_usd, generation_spend_cap_usd on Formats, and a balance check.
- Captions out of sync with the audio: check STT word times and offsets
Captions running early or late usually trace to an unapplied offset. How Sume STT word times work, which offset to add, and a Python merge that applies it.
- Connect a new MCP client to Sume: five calls that prove it works
After you add https://mcp.sume.com/mcp to a new client, run mcp_health, tools_list, tools_schema, account_me and catalog_list. What each result should show.
Written by Sume