Inworld buys Ultravox, voices move to TTS-2: pin voice and model IDs
Inworld's Ultravox deal moves built-in voices to Realtime TTS-2 at no extra cost. Why to pin voice and model ids wherever you generate speech.

After Inworld's acquisition of Ultravox, announced on September 30, built-in Inworld voices on Ultravox move to Realtime TTS-2 at no extra cost, and existing voice IDs keep working. The blog also says that if you use a voice from another provider, nothing changes today. The practical lesson for any speech pipeline is to pin the voice and model you tested, and to log them with every output.
A silent engine swap can change timbre, pacing and pronunciation without breaking a single API call. If your brand voice is part of a product, that matters more than a small price move.
What the two Inworld posts state
Both rows below come from Inworld's own pages.
| Item | What the page says |
|---|---|
| Acquisition post | Sep 30: Inworld acquires Ultravox |
| Built-in Inworld voices on Ultravox | Move to Realtime TTS-2 at no extra cost |
| Existing voice IDs | Keep working |
| Voices from another provider | Nothing changes today |
| Realtime TTS-2 | GA Aug 31; 100+ languages; sub-200 ms median time to first audio |
| Voice cloning | From 5 to 15 seconds of audio |
| Billing | Metered per character |
How to guard against silent voice changes
- Store the voice id, model id and request date next to each generated file.
- Keep one approved reference clip per voice and compare a fresh take against it before each release.
- Alert on a change in job.model or job.request values in your logs.
- Do not use a floating alias such as sonic-latest for finished work.
Pinning on Sume
On the Sume TTS Router you pass a catalog model such as sonic-3.6 and a voice.id or avatar handle; job.model echoes the model you requested. The alias sonic-latest resolves to sonic-3.6 today, so a pinned id is safer than the alias; see pin the TTS model id.
import os, requests
r = requests.post(
"https://api.sume.com/v1/tts-router/generate",
headers={
"Authorization": f"Bearer {os.environ['SUME_API_KEY']}",
"Idempotency-Key": "tts-demo-001",
},
json={
"model": "sonic-3.6",
"transcript": "Welcome back. Today we compare three prices.",
"voice": {"id": os.environ["SUME_VOICE_ID"]},
"timestamps": {"words": True},
},
timeout=30,
)
r.raise_for_status()
print(r.json())What a vendor change can break
Pricing and access rarely break a voice pipeline. Quality drift does. A new engine can read numbers differently, add a breath where the old one did not, or shift the speaking rate by a few percent. For a one-minute voiceover that is a minor edit; for a catalogue of 500 product videos it is a re-approval cycle.
Inworld says voice IDs keep working through the move, which is the right promise for compatibility. It does not promise that the audio will sound identical, and the post does not claim that. Plan a listening pass on your three most important scripts the week a vendor announces a change, and keep the results in a shared folder.
What Sume does not do
Sume does not offer Inworld voices or models. Sume TTS covers Cartesia Sonic only. The router has no Eleven, OpenAI, Gemini, MAI or Inworld engines, no streaming TTS, and no routing presets. Jobs are asynchronous, and any audio over 1,200 seconds fails with tts_duration_exceeded. I did not verify how Ultravox prices voices, only what the two Inworld pages state.
Takeaway
Treat a voice as a versioned asset. When a vendor merger or a model refresh happens, you will know in minutes whether your output changed, because you have the old take to compare. See Cartesia model sunsets on October 20 for another change that rewards pinned ids.
Sources
Related posts
More in Models
- Is Hy Image 3.5 Preview on Sume? Check the models list in code
Tencent's Hy Image 3.5 Preview is not in the Sume rows used here. How to confirm against GET /v1/images/models with a 12-line script instead of guessing an id.
- Is MAI-Voice-2.1 on Sume? What Sume's audio docs list instead
Sume does not list MAI-Voice-2.1 or MAI-Transcribe-2. Here is what its docs do list for speech, transcripts, music and captions, with prices.
- Kling 3 audio on vs off for 100 fifteen-second SKU ads: $105 gap
Kling 3 Pro on Sume is $0.14 per second with audio off and $0.21 on. At 15 s across 100 SKUs that is $210 vs $315, a $105 difference.
- Kling 3 stops at 15 seconds: three ways to a 30-second clip
Kling 3 on Sume takes 4 to 15 seconds. For 30 seconds, use Wan 3.0 at $3.75 (720p) or Seedance 2.5, or join two takes with Timeline. Prices side by side.
Written by Sume