Inworld TTS 2 style steering and the Sume voiceover path

What is reported about Inworld TTS 2 style steering and 100+ languages, and how to produce a voiceover with Sume's tts_create tool and join takes.

3 min readSume
All posts

Inworld TTS 2 is reported to support style steering across more than 100 languages. On Sume you create a voiceover with the hosted MCP tool tts_create, then join the takes into one gapless file with Timeline audio. The Sume docs do not list TTS model ids, so read the live tool schema instead of assuming a name.

What is reported

A Versely roundup of the voice market lists these dates and claims. Treat them as that article's reporting, not as Sume facts.

Inworld releases as reported by Versely (read 2026-10-03)
ReleaseDateReported detail
Inworld TTS 2May 5, 2026100+ languages, style steering, Research Preview
Sonic 3.5June 16, 2026Listed in the same roundup

What this post does not claim

The roundup does not give prompt wording for style steering in the part this post draws on, so there is no list of prompts here. Check Inworld's own documentation for how steering is written before you rely on it.

The Sume side

Sume exposes speech synthesis on hosted MCP as tts_create. It is a paid tool, so it needs idempotency_key, and under OAuth it needs mcp:write. Two free read tools help with scripts: tts_source_get returns the accepted-script manifest for a tts_create that uses transcript_source, and tts_source_verify_spine checks selected TTS jobs against that script.

Because the docs do not name TTS model ids, call tools_schema for tts_create in your own session and read the contract it returns. That is the source of truth for voices and options.

Joining takes

Generate one take per sentence or paragraph, then join them. POST /v1/timeline-1.0/audio with operation: "concat" takes up to 20 ordered parts[] and joins them in the sample domain, so there is no re-synthesis and no silence at the seams. The result gives segments[] with start offsets that you can use to place video slots. The job is a flat $0.01 per job, and you should confirm that live in GET /v1/catalog.

Keep wav, the default, when the file will be joined again or will drive lip-sync. mp3 is smaller but re-adds priming padding at every edge.

Lip-sync note

Sume's models doc is explicit that video models do not lip-sync to generated TTS or to a later voice-over. A talking face is a Fabric clip built from an accepted still plus the TTS audio. If your goal is a voiced on-camera shot, plan for Fabric rather than narration laid under a video-model clip.

Sources

Related posts

More in Models

All Models posts

Written by Sume