Voice drift in long narration: Nova 2 Sonic's 52% vs Sume chunks

Speaker drift is what splits a long voiceover into jobs. Amazon claims -52% on an internal set; here is how to measure drift across Sume TTS chunks yourself.

5 min readSume
All posts

When a long narration is cut into several TTS jobs, the risk is that chunk four sounds like a different person than chunk one. Amazon's May 2026 Nova 2 Sonic notes claim 52% less speaker drift, on an internal data set. On Sume the practical defence is simple: use the same voice id and the same model id for every chunk, keep each chunk at or under 20,000 characters, and compare the chunks before you stitch them.

Where drift comes from in a chunked job

Sume does not carry state from one TTS job to the next. Each job is an independent request, so the only things that tie chunks together are the inputs you repeat: the voice selector, the router model id, the language, and generation_config values such as speed. Change any of them mid-script and you can hear it. The most common accident is a model alias that moves: sonic-latest resolves to sonic-3.6 today, so pin an explicit id for a long project instead.

Chunk boundaries matter too. Split at sentence ends, never mid-sentence, and keep the sentence ids stable if you use transcript_source, because the receipt lists them in order.

Settings to hold constant across chunks, Sume TTS router docs and Amazon release notes, read 2026-10-05
SettingKeep fixedWhy
modelOne explicit id such as sonic-3.6Aliases can move; pass-through catalog rows are named
voice.id or avatar_handleSame value in every chunkVoice is the main identity signal
generation_config.speed / volumeSame valuesSpeed 0.6-1.5 and volume 0.5-2 are per-job
languageSame codeMismatch returns 409 tts_voice_language_mismatch
Chunk sizeUp to 20,000 charactersLongest accepted transcript per job

Measure it before you ship

Do not trust your ears on the first listen. Generate the same 300-character calibration sentence at the top of every chunk, or as a separate 1-cent job per chunk, and compare durations. Two renders of identical text with the same voice should land within a small spread; if one chunk's calibration line is noticeably longer or shorter, its pacing drifted, and you can regenerate that one chunk alone. A 210-character calibration line costs 1 cent, a 211-character line costs 2, so keep it short.

Amazon's -52% figure tells you drift is a known failure of long generations. It cannot tell you how your voice behaves, which is why the calibration job is worth the cent.

Stitching safely

Request wav output for chunks you will join. Concatenating mp3 files adds frame padding at each seam, while wav concatenates cleanly. Use segmentation with timestamps.words when you need per-sentence slices for captions. A job over 1,200 seconds of audio fails with tts_duration_exceeded and is not captured, so a very slow speed setting can push a chunk over the limit even when the text is under 20,000 characters.

What to log per chunk

Keep a small ledger: chunk number, job id, model id echoed in the job, voice id, character count and the cents billed. If a listener flags chunk four, the ledger tells you whether any setting differed from chunk three. It also lets you re-run only the bad chunk, and an Idempotency-Key per chunk stops a retry from double-billing the same text.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume