AssemblyAI text to speech: coming soon, and what Sume has today

AssemblyAI's product menu lists a Text-to-Speech API as coming soon. Sume TTS is available now: what a call takes, the 20000-character cap, and word timings.

4 min readSume
All posts

AssemblyAI's product menu, read 2026-10-01, shows "Text-to-Speech API" with a "coming soon" label beside the Voice Agent API and Dictation API. Sume TTS is a live endpoint today: POST /v1/tts-1.0/generate, which takes a transcript and returns audio.

This post says nothing about AssemblyAI's unreleased product beyond that label.

What does the AssemblyAI page say?

Only that a Text-to-Speech API is coming soon. No pricing, voices or dates appear in the snapshot, so there is nothing to compare.

What does a Sume TTS call take?

The schema calls it the Sume TTS 1.0 text-to-speech request. Send a transcript (maximum 20000 characters) and pick a voice with avatar_id / avatar_handle or voice.id. Optional timestamps.words and segmentation.mode=sentence return word timings.

TTS routes, read 2026-10-01.
RouteUse
POST /v1/tts-1.0/generateSume TTS 1.0
POST /v1/tts-router/generateA named engine from the Router catalog
GET /v1/tts-router/modelsList Router models

When would I use the Router route?

The Router request requires a TTS Router catalog model id, so you name the engine. Use TTS 1.0 when you do not care which engine runs.

What do I do with the audio afterwards?

Word timings feed captions or cut lists, and for join and slice work see mp3 or wav. The overview is in text to speech API.

Sources

Related posts

More in Models

All Models posts

Written by Sume