Decagon Voice 3 handles 70+ languages mid-call. A TTS job takes one
Voice 3 detects language and switches mid-sentence. A Sume TTS job speaks one language per request. How to plan a multilingual script around that.

Decagon says Voice 3 supports 70+ languages with automatic language detection and can switch languages in the middle of a sentence. A Sume text-to-speech job is the opposite shape: one transcript, one voice and one target language per request, so a bilingual script becomes several jobs.
What Decagon claims for Voice 3
From the October 1, 2026 announcement:
- 70+ languages with automatic language detection.
- Mid-sentence language switching, so a caller can mix languages.
- Locale-specific voices validated by native speakers.
How a Sume TTS job handles language
tts_create takes a language in the payload along with a voice selector. Sume checks the voice's recorded language against the request. If they clearly disagree, the API answers 409 tts_voice_language_mismatch before it creates a job or charges anything, and the MCP tool shows a warning that needs the user to confirm.
That guard matters for multilingual work. A Spanish script sent to an English-recorded voice is a real mistake that costs money to find by ear. Retry only after a person confirms, using confirm_language_mismatch.
Splitting a mixed-language script
A line like a brand slogan in English inside a Spanish ad is a common case. The practical way to handle it with a job-based tool is to cut the script at the language boundary and make each part its own job, then join the audio with timeline audio concat, which joins up to 20 parts into one gapless file for $0.01 flat per job.
Compare the plan with a live agent. Voice 3 decides the language as the caller speaks. A job fixes it when you submit.
| Need | Voice 3 (per Decagon) | Sume TTS job |
|---|---|---|
| Detect the language | Automatic | You set it |
| Switch mid-sentence | Yes | Split into jobs, then concat |
| Wrong voice for language | Not described | 409 before any charge, confirm to override |
| Join pieces | Not applicable | Timeline audio concat, up to 20 parts |
Budget a split script
Cost follows characters, not languages. At $0.0475 per 1,000 characters, a 600-character Spanish line and a 40-character English slogan are two jobs, each rounded up to a cent, plus $0.01 to join them. Keep the pieces at sentence boundaries so the concat has no audible seam.
For one voice across languages, check the voice's language tags first and keep the same voice id when the voice supports the target language.
Sources
Related posts
More in Comparisons
- Decagon Voice 3 duplex agent or a voiceover job: which do you need?
Voice 3 listens while it speaks. A voiceover job does not. How to choose between a live voice agent and a file-based TTS job for your project.
- Does Sume have a real-time avatar API? No, here is what it has instead
Sume has no live avatar session. It has async avatar jobs: create an avatar, render a talking video, or lip-sync a still to audio. Routes, limits and prices.
- Edits on desktop or an API render: which for a weekly Reel?
Edits now has a desktop app and an AI assistant. Use it for creative one-offs and a Sume render for repeated, logged Reels: a decision table with job prices.
- Eleven v4 Turbo in ElevenAgents vs Sume TTS as async jobs
Eleven v4 Turbo targets live agents. Sume TTS is an async job with poll or webhook and no streaming. Which one fits a call bot and which fits produced audio.
Written by Sume