Voice drift in long narration: Nova 2 Sonic's 52% vs Sume chunks
Speaker drift is what splits a long voiceover into jobs. Amazon claims -52% on an internal set; here is how to measure drift across Sume TTS chunks yourself.

When a long narration is cut into several TTS jobs, the risk is that chunk four sounds like a different person than chunk one. Amazon's May 2026 Nova 2 Sonic notes claim 52% less speaker drift, on an internal data set. On Sume the practical defence is simple: use the same voice id and the same model id for every chunk, keep each chunk at or under 20,000 characters, and compare the chunks before you stitch them.
Where drift comes from in a chunked job
Sume does not carry state from one TTS job to the next. Each job is an independent request, so the only things that tie chunks together are the inputs you repeat: the voice selector, the router model id, the language, and generation_config values such as speed. Change any of them mid-script and you can hear it. The most common accident is a model alias that moves: sonic-latest resolves to sonic-3.6 today, so pin an explicit id for a long project instead.
Chunk boundaries matter too. Split at sentence ends, never mid-sentence, and keep the sentence ids stable if you use transcript_source, because the receipt lists them in order.
| Setting | Keep fixed | Why |
|---|---|---|
| model | One explicit id such as sonic-3.6 | Aliases can move; pass-through catalog rows are named |
| voice.id or avatar_handle | Same value in every chunk | Voice is the main identity signal |
| generation_config.speed / volume | Same values | Speed 0.6-1.5 and volume 0.5-2 are per-job |
| language | Same code | Mismatch returns 409 tts_voice_language_mismatch |
| Chunk size | Up to 20,000 characters | Longest accepted transcript per job |
Measure it before you ship
Do not trust your ears on the first listen. Generate the same 300-character calibration sentence at the top of every chunk, or as a separate 1-cent job per chunk, and compare durations. Two renders of identical text with the same voice should land within a small spread; if one chunk's calibration line is noticeably longer or shorter, its pacing drifted, and you can regenerate that one chunk alone. A 210-character calibration line costs 1 cent, a 211-character line costs 2, so keep it short.
Amazon's -52% figure tells you drift is a known failure of long generations. It cannot tell you how your voice behaves, which is why the calibration job is worth the cent.
Stitching safely
Request wav output for chunks you will join. Concatenating mp3 files adds frame padding at each seam, while wav concatenates cleanly. Use segmentation with timestamps.words when you need per-sentence slices for captions. A job over 1,200 seconds of audio fails with tts_duration_exceeded and is not captured, so a very slow speed setting can push a chunk over the limit even when the text is under 20,000 characters.
What to log per chunk
Keep a small ledger: chunk number, job id, model id echoed in the job, voice id, character count and the cents billed. If a listener flags chunk four, the ledger tells you whether any setting differed from chunk three. It also lets you re-run only the bad chunk, and an Idempotency-Key per chunk stops a retry from double-billing the same text.
Sources
Related posts
More in Use cases
- Luggage holiday travel ad: fit a 16:9 clip to 9:16 with blur for $0.11
Reuse a 16:9 luggage ad as a 9:16 Reel: detach the audio for $0.01, then one Timeline render with fit blur for $0.10, instead of $1.00 to regenerate.
- Live transcription for Korean meetings: MAI-Transcribe-2-Streaming
MAI-Transcribe-2-Streaming claims 60 languages with auto detection and about 100ms to first text. Verify Korean yourself. Sume transcribes files, not live.
- Make an AI voice read an email or order code right: spell, verify
Amazon and Cartesia both claim better codes, emails and phone numbers. The reliable fix is text prep plus a read-back check. Try both on Sume's TTS router.
- Mattress Black Friday video ad: a dimmed bedroom clip and a price card
Make a mattress Black Friday ad with AI: one 8-second bedroom clip, a dim pass, and a price-card overlay on top, about $1.12 in Sume jobs.
Written by Sume