Nova 2.5 Sonic for a narration file? Speech-to-speech vs Sume TTS
Nova 2.5 Sonic is built for live voice agents. For a finished narration file, Sume TTS takes a transcript up to 20,000 characters and returns audio as a job.

If you need a narration file, a voiceover for a video or a read-aloud track, you do not need Nova 2.5 Sonic. AWS describes it as a speech-to-speech model for real-time voice agents, where a caller speaks and the model answers by voice in a live session. For a finished file, submit a transcript to Sume TTS as an asynchronous job and download the audio when it completes.
The Nova facts are from AWS's October 2026 announcement (read 2026-10-11): general availability on October 5, 2026 in four regions, a 256K context window, expressive voices in seven languages, voice and text in the same session, and the same pricing as Nova 2 Sonic. AWS lists customer service, personal assistants and real-time conversational applications as the primary uses. None of that is about producing an offline audio file.
What is the difference between the two jobs?
A live agent decides what to say as it listens, and its latency is part of the product. A narration job has the words already, and its quality is judged on the file. That changes the interface you want.
| Question | Nova 2.5 Sonic | Sume TTS 1.0 |
|---|---|---|
| Who writes the words | The model, during the conversation | You, as a transcript |
| Interface | Real-time speech session | Async job: submit, poll, read result |
| Input limit | 256K context window | 1 to 20,000 characters per request |
| Price shape | Same pricing as Nova 2 Sonic | $0.0475 per 1,000 characters, at most $0.95 for a full request |
| Output | Live audio in the session | Audio file with optional word timestamps |
How does a Sume narration job work?
POST /v1/tts-1.0/generate takes a transcript and a voice selector, either an avatar_id or avatar_handle or a voice id from your library. You can set language, output format, timestamps and segmentation. The job is asynchronous: read the status, then the result. If the voice's stored language differs from the requested one, the API answers 409 tts_voice_language_mismatch before it creates a job or charges anything, so you can confirm and retry.
If you want to pick the engine yourself, POST /v1/tts-router/generate takes a required model from the catalog, such as sonic-3.6, sonic-3.5 or sonic-latest. The router bills character-metered list prices times 1.25.
How do you cost a script before you submit?
Count the characters of the transcript, not the words or the minutes. At $0.0475 per 1,000 characters, a full 20,000-character request tops out at $0.95, which the catalog states as the maximum, and the minimum charge is one cent. Over MCP a dry run previews the cost without spending, and a max_spend_usd cap can be set on a paid create.
For anything longer than 20,000 characters, split the script at sentence boundaries into several requests and join the audio afterward (a request over 1,200 seconds of audio fails with tts_duration_exceeded), for example with the timeline audio route. Splitting at sentences keeps the pauses natural and lets you regenerate one slice without redoing the rest.
When would you combine the two?
A voice agent can still hand off to a file job. Suppose a caller approves a script during a call. The agent's tool can submit the approved text to Sume TTS and return a job id, as in an async tool call, and the finished file can then be mixed into a video. The agent talks; Sume makes the deliverable.
If you record a call and need its text, that is the opposite direction: Sume STT turns audio into words at $0.01 per audio minute.
Sources
Related posts
More in Models
- Qwen-Image-2.1-Turbo runs 8 steps; what does Sume's Qwen row expose?
Steps, CFG and seed are model-card settings. Sume's qwen/qwen-image row lists none of them: only ratio, n 1-4, references, output format. Check with one GET.
- Reka Rho-1: can you generate video with it today?
Reka Rho-1 is a 19B research preview announced October 5, 2026, with no public weights, API or demo link. For video you can run today, check Sume's catalog.
- Grok Imagine video ids: three at xAI, one on Sume
xAI lists grok-imagine-video, 1.5 and 1.5-lite; Sume's catalog has only grok-imagine-video-1.5. A quick map so you send the id that exists.
- What is Grok Imagine Video 1.5 Lite, and what does Sume run?
Grok Imagine Video 1.5 Lite is xAI's low-price video model at $0.02 a second. Sume's catalog has grok-imagine-video-1.5 instead, image-to-video only.
Written by Sume