Decagon Voice 3 handles 70+ languages mid-call. A TTS job takes one

Voice 3 detects language and switches mid-sentence. A Sume TTS job speaks one language per request. How to plan a multilingual script around that.

5 min readSume
All posts

Decagon says Voice 3 supports 70+ languages with automatic language detection and can switch languages in the middle of a sentence. A Sume text-to-speech job is the opposite shape: one transcript, one voice and one target language per request, so a bilingual script becomes several jobs.

What Decagon claims for Voice 3

From the October 1, 2026 announcement:

  • 70+ languages with automatic language detection.
  • Mid-sentence language switching, so a caller can mix languages.
  • Locale-specific voices validated by native speakers.

How a Sume TTS job handles language

tts_create takes a language in the payload along with a voice selector. Sume checks the voice's recorded language against the request. If they clearly disagree, the API answers 409 tts_voice_language_mismatch before it creates a job or charges anything, and the MCP tool shows a warning that needs the user to confirm.

That guard matters for multilingual work. A Spanish script sent to an English-recorded voice is a real mistake that costs money to find by ear. Retry only after a person confirms, using confirm_language_mismatch.

Splitting a mixed-language script

A line like a brand slogan in English inside a Spanish ad is a common case. The practical way to handle it with a job-based tool is to cut the script at the language boundary and make each part its own job, then join the audio with timeline audio concat, which joins up to 20 parts into one gapless file for $0.01 flat per job.

Compare the plan with a live agent. Voice 3 decides the language as the caller speaks. A job fixes it when you submit.

Language handling compared (read 2026-10-07)
NeedVoice 3 (per Decagon)Sume TTS job
Detect the languageAutomaticYou set it
Switch mid-sentenceYesSplit into jobs, then concat
Wrong voice for languageNot described409 before any charge, confirm to override
Join piecesNot applicableTimeline audio concat, up to 20 parts

Budget a split script

Cost follows characters, not languages. At $0.0475 per 1,000 characters, a 600-character Spanish line and a 40-character English slogan are two jobs, each rounded up to a cent, plus $0.01 to join them. Keep the pieces at sentence boundaries so the concat has no audible seam.

For one voice across languages, check the voice's language tags first and keep the same voice id when the voice supports the target language.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume