Inworld TTS 2 style steering and the Sume voiceover path
What is reported about Inworld TTS 2 style steering and 100+ languages, and how to produce a voiceover with Sume's tts_create tool and join takes.

Inworld TTS 2 is reported to support style steering across more than 100 languages. On Sume you create a voiceover with the hosted MCP tool tts_create, then join the takes into one gapless file with Timeline audio. The Sume docs do not list TTS model ids, so read the live tool schema instead of assuming a name.
What is reported
A Versely roundup of the voice market lists these dates and claims. Treat them as that article's reporting, not as Sume facts.
| Release | Date | Reported detail |
|---|---|---|
| Inworld TTS 2 | May 5, 2026 | 100+ languages, style steering, Research Preview |
| Sonic 3.5 | June 16, 2026 | Listed in the same roundup |
What this post does not claim
The roundup does not give prompt wording for style steering in the part this post draws on, so there is no list of prompts here. Check Inworld's own documentation for how steering is written before you rely on it.
The Sume side
Sume exposes speech synthesis on hosted MCP as tts_create. It is a paid tool, so it needs idempotency_key, and under OAuth it needs mcp:write. Two free read tools help with scripts: tts_source_get returns the accepted-script manifest for a tts_create that uses transcript_source, and tts_source_verify_spine checks selected TTS jobs against that script.
Because the docs do not name TTS model ids, call tools_schema for tts_create in your own session and read the contract it returns. That is the source of truth for voices and options.
Joining takes
Generate one take per sentence or paragraph, then join them. POST /v1/timeline-1.0/audio with operation: "concat" takes up to 20 ordered parts[] and joins them in the sample domain, so there is no re-synthesis and no silence at the seams. The result gives segments[] with start offsets that you can use to place video slots. The job is a flat $0.01 per job, and you should confirm that live in GET /v1/catalog.
Keep wav, the default, when the file will be joined again or will drive lip-sync. mp3 is smaller but re-adds priming padding at every edge.
Lip-sync note
Sume's models doc is explicit that video models do not lip-sync to generated TTS or to a later voice-over. A talking face is a Fabric clip built from an accepted still plus the TTS audio. If your goal is a voiced on-camera shot, plan for Fabric rather than narration laid under a video-model clip.
Sources
Related posts
More in Models
- Nano Banana 3: does it exist? Current ids
Google autocomplete suggests a Nano Banana 3, but suggestions are not releases. What autocomplete lists and which Nano Banana id Sume accepts.
- Lip sync with a hand over the mouth: sync-3 vs Fabric on Sume
Sync Labs says sync-3 handles obstructions on faces. Sume's lip-sync routes start from a still plus audio. What each takes as input and what is promised.
- Luma's 2026 timeline: Ray3.14, Ray3.2, Scenes and Variants
Luma shipped Ray3.14 in January, Ray3.2 in June, Scenes in August and Variants on Oct 1, 2026. What each added, and why to pin model ids.
- Lyria 3.5 blocks artist-voice prompts: how to write briefs that pass
Google's Lyria 3.5 docs note that prompts asking for specific artist voices are blocked. Describe the sound instead, then run it through the Sume Music Router.
Written by Sume