AssemblyAI text to speech: coming soon, and what Sume has today
AssemblyAI's product menu lists a Text-to-Speech API as coming soon. Sume TTS is available now: what a call takes, the 20000-character cap, and word timings.

AssemblyAI's product menu, read 2026-10-01, shows "Text-to-Speech API" with a "coming soon" label beside the Voice Agent API and Dictation API. Sume TTS is a live endpoint today: POST /v1/tts-1.0/generate, which takes a transcript and returns audio.
This post says nothing about AssemblyAI's unreleased product beyond that label.
What does the AssemblyAI page say?
Only that a Text-to-Speech API is coming soon. No pricing, voices or dates appear in the snapshot, so there is nothing to compare.
What does a Sume TTS call take?
The schema calls it the Sume TTS 1.0 text-to-speech request. Send a transcript (maximum 20000 characters) and pick a voice with avatar_id / avatar_handle or voice.id. Optional timestamps.words and segmentation.mode=sentence return word timings.
| Route | Use |
|---|---|
POST /v1/tts-1.0/generate | Sume TTS 1.0 |
POST /v1/tts-router/generate | A named engine from the Router catalog |
GET /v1/tts-router/models | List Router models |
When would I use the Router route?
The Router request requires a TTS Router catalog model id, so you name the engine. Use TTS 1.0 when you do not care which engine runs.
What do I do with the audio afterwards?
Word timings feed captions or cut lists, and for join and slice work see mp3 or wav. The overview is in text to speech API.
Sources
Related posts
More in Models
- Creatify Boreal talking clips vs Sume's still-plus-audio route
Creatify says Boreal's gains are smallest on single-person talking clips. Sume makes every speaking shot from an accepted still plus TTS audio via Fabric.
- Deepgram nova-3-pharma vs Sume STT: drug-name transcripts
Deepgram added nova-3-pharma for English drug names. Sume STT has one public model, sume/stt-1.0, so check each drug name against word timings.
- ElevenLabs character limits by model vs Sume TTS 20,000
ElevenLabs lists 5,000 characters for v3, 10,000 for v4 and 40,000 for Flash v2.5. Sume TTS 1.0 takes up to 20,000 characters in one request.
- ElevenLabs speech to speech API: what Sume offers instead
ElevenLabs lists speech-to-speech voice changer models. Sume has no voice-to-voice endpoint: transcribe with STT, edit the text, then run TTS.
Written by Sume