Voice model updated in place with no API change: how to detect it
Nova 2 Sonic was refreshed in place in May with no API change. If a vendor can change your voice silently, log the model id and a canary clip. Sume code inside.

Amazon's release notes say the May 2026 Nova 2 Sonic refresh was deployed in place over May 21-28 with no API change. That is good for quality and bad for reproducibility: your integration did not change, but the audio might have. The defence is a canary. Keep one fixed sentence, render it on a schedule, store the model id the API echoes back, and alert when the duration moves. Sume echoes the requested model id in job.model, so you can log it with every job.
Two kinds of change
A vendor can change audio in two ways. A named version bumps, like Cartesia publishing sonic-3.6 and a dated snapshot sonic-3.6-2026-08-27 on its changelog. Or the behaviour behind a stable name shifts without a new name, which is what an in-place refresh is. The first you can pin. The second you can only detect.
On the Sume TTS router, the catalog rows are sonic-3.6, sonic-3.5, sonic-3, sonic-latest and sonic-preview. sonic-latest is an alias for sonic-3.6, and sonic-preview is a beta channel that can change without notice. An explicit older row such as sonic-3.5 keeps a project on the version it was approved on.
| Change | Example from the pages read | What catches it |
|---|---|---|
| Named version bump | Sonic 3.6 released Aug 27 | Pin the router model id |
| In-place refresh | Nova 2 Sonic May 21-28, no API change | Canary clip and duration alert |
| Alias move | sonic-latest resolves to sonic-3.6 | Pin explicit id; log job.model |
| Beta channel | sonic-preview can change without notice | Do not use in production |
A canary in 15 lines
Render the same 120-character sentence daily, and keep the job's reported usage. The sentence costs 1 cent (120 x 0.00475 = 0.57 cents, rounded up). Store the model id and the audio length, and compare against the first approved render. Seconds-level drift in length is a coarse signal but it is free to compute.
import os, time, requests
H = {"x-api-key": os.environ["SUME_API_KEY"]}
B = "https://api.sume.com"
def run(path, body):
d = requests.post(B + path, json=body, headers=H, timeout=60).json()["data"]
while not d.get("terminal"):
time.sleep(d.get("next_poll_after_seconds") or 2)
d = requests.get(d["status_url"], headers=H, timeout=60).json()["data"]
return requests.get(d["result_url"], headers=H, timeout=60).json()
line = "Your appointment is confirmed for Tuesday at 4 PM."
res = run("/v1/tts-router/generate", {"model": "sonic-3.5", "transcript": line,
"voice": {"id": os.environ["SUME_VOICE_ID"]}})
open("canary.json", "w").write(__import__("json").dumps(res, indent=2))
print(res) # log job.model and the audio duration with a dateWhat a canary cannot do
A length check will not catch a subtle change in timbre. Keep the audio file itself for a human spot check each month, and re-listen when your log shows the vendor announced a refresh. Treat the canary as the alarm, not the judge.
How often to run the canary
Daily is enough for most teams, and it costs 1 cent a day, about 30 cents a month. Run it right before any large batch as well, because a regression found before 400 jobs is cheaper than one found after. Keep the last 30 results so a slow trend is visible, not just a jump.
Also record the vendor's own notice dates. The Amazon page lists the in-place deployment window as May 21-28, so a canary that moved on May 24 would have matched a documented event rather than a mystery. When your alarm has no matching notice, treat it as an incident and ask the vendor.
What to do when the canary moves
First re-run it once, since a single outlier is common. If it repeats, listen to both files back to back. If the change is audible and unwanted, move the project to an older explicit router row such as sonic-3.5 while you investigate. If it is inaudible, update the baseline and write down why. Either way the log gives you a date to quote when you ask the vendor what changed.
Sources
Related posts
More in Developers
- spend_approval_queue_full 429: clear pending approvals, do not retry
A thread with too many pending spend approvals gets 429 spend_approval_queue_full. Resolve the pending ones first; a 503 store_misconfigured is for support.
- spend_confirmation_required 402 on Sume: why retrying will not help
A 402 spend_confirmation_required means a person must approve the spend first. It is not a balance error: retryable is false and next_action is fix_input.
- Split a 10-minute TikTok into parts for 3 and 5-minute accounts
TikTok's API allows up to 10 minutes, but an account may be limited to 3 or 5. Split a 600-second video into 4 or 2 parts with Sume trim at $0.02 a job.
- Split a 40-second brief into four Omni prompts, one character block
A Python script that turns one 40-second brief into four 10-second Gemini Omni 1.1 Flash request bodies for Sume, with one shared character block.
Written by Sume