Nova 2.5 Sonic for a narration file? Speech-to-speech vs Sume TTS

Nova 2.5 Sonic is built for live voice agents. For a finished narration file, Sume TTS takes a transcript up to 20,000 characters and returns audio as a job.

4 min readSume
All posts

If you need a narration file, a voiceover for a video or a read-aloud track, you do not need Nova 2.5 Sonic. AWS describes it as a speech-to-speech model for real-time voice agents, where a caller speaks and the model answers by voice in a live session. For a finished file, submit a transcript to Sume TTS as an asynchronous job and download the audio when it completes.

The Nova facts are from AWS's October 2026 announcement (read 2026-10-11): general availability on October 5, 2026 in four regions, a 256K context window, expressive voices in seven languages, voice and text in the same session, and the same pricing as Nova 2 Sonic. AWS lists customer service, personal assistants and real-time conversational applications as the primary uses. None of that is about producing an offline audio file.

What is the difference between the two jobs?

A live agent decides what to say as it listens, and its latency is part of the product. A narration job has the words already, and its quality is judged on the file. That changes the interface you want.

Live voice agent vs narration file (AWS announcement and Sume repo, read 2026-10-11)
QuestionNova 2.5 SonicSume TTS 1.0
Who writes the wordsThe model, during the conversationYou, as a transcript
InterfaceReal-time speech sessionAsync job: submit, poll, read result
Input limit256K context window1 to 20,000 characters per request
Price shapeSame pricing as Nova 2 Sonic$0.0475 per 1,000 characters, at most $0.95 for a full request
OutputLive audio in the sessionAudio file with optional word timestamps

How does a Sume narration job work?

POST /v1/tts-1.0/generate takes a transcript and a voice selector, either an avatar_id or avatar_handle or a voice id from your library. You can set language, output format, timestamps and segmentation. The job is asynchronous: read the status, then the result. If the voice's stored language differs from the requested one, the API answers 409 tts_voice_language_mismatch before it creates a job or charges anything, so you can confirm and retry.

If you want to pick the engine yourself, POST /v1/tts-router/generate takes a required model from the catalog, such as sonic-3.6, sonic-3.5 or sonic-latest. The router bills character-metered list prices times 1.25.

How do you cost a script before you submit?

Count the characters of the transcript, not the words or the minutes. At $0.0475 per 1,000 characters, a full 20,000-character request tops out at $0.95, which the catalog states as the maximum, and the minimum charge is one cent. Over MCP a dry run previews the cost without spending, and a max_spend_usd cap can be set on a paid create.

For anything longer than 20,000 characters, split the script at sentence boundaries into several requests and join the audio afterward (a request over 1,200 seconds of audio fails with tts_duration_exceeded), for example with the timeline audio route. Splitting at sentences keeps the pauses natural and lets you regenerate one slice without redoing the rest.

When would you combine the two?

A voice agent can still hand off to a file job. Suppose a caller approves a script during a call. The agent's tool can submit the approved text to Sume TTS and return a job id, as in an async tool call, and the finished file can then be mixed into a video. The agent talks; Sume makes the deliverable.

If you record a call and need its text, that is the opposite direction: Sume STT turns audio into words at $0.01 per audio minute.

Sources

Related posts

More in Models

All Models posts

Written by Sume