Nova 2 Sonic runs in 4 AWS regions; can you pick one on Sume TTS?
Amazon lists Nova 2 Sonic GA in N. Virginia, Oregon, Tokyo and Stockholm. Sume's TTS has one base URL and no region field; here is what that means for you.

No: on Sume you cannot pick a region for text-to-speech. The request schema has no region field, and every call goes to one base URL, https://api.sume.com. Amazon's Nova 2 release notes list general availability for Nova 2 Sonic in four AWS regions: US East (N. Virginia), US West (Oregon), Tokyo and Stockholm. If your rule says audio must be generated in a named region, that is a fact to settle before you choose an API, and a job-based service with one endpoint will not satisfy it.
What the Amazon page lists
The same release notes describe two refreshes. A March 2026 refresh added Polly-compatible voices, cut p50 latency by 150 ms and improved turn-taking on 8 kHz telephony audio. A May 2026 refresh reported 88% fewer hallucinations, 52% less speaker drift and 28% fewer critical errors, each on an internal data set, and was deployed in place between May 21 and 28 with no API change. Regions matter because a real-time voice agent wants the model close to the caller.
| Question | Amazon release notes | Sume TTS 1.0 / router |
|---|---|---|
| Where does it run? | US East (N. Virginia), US West (Oregon), Tokyo, Stockholm | One base URL, https://api.sume.com; no region field |
| Telephony audio | March 2026: better 8 kHz turn-taking | output_format sample_rate can be 8000 |
| Model selection | Model refreshed in place | Router model id chosen per request |
| Delivery | Live sessions | Async job; sync wait up to 30 s; webhook |
When a region matters, and when it does not
Region matters for two reasons: latency for live speech, and data residency rules. Prepared audio, such as a narrated video, is not latency-bound, because nobody waits for the first byte. Data residency is a policy question, and the honest answer for Sume is that a request is handled through its API without a region choice, so if your policy names a region, check with your own compliance owner before you send text containing personal data.
Sume does give you ways to limit what is sent. The metadata object you attach to a TTS job is stored with the job and is not sent to the provider, so keep customer ids there instead of in the transcript.
import os, requests
r = requests.post("https://api.sume.com/v1/tts-router/generate",
headers={"x-api-key": os.environ["SUME_API_KEY"]},
json={"model": "sonic-3.6",
"transcript": "Your appointment is at 4 PM on Tuesday.",
"voice": {"id": os.environ["SUME_VOICE_ID"]},
"output_format": {"container": "wav", "sample_rate": 8000,
"encoding": "pcm_mulaw"},
"metadata": {"customer_ref": "c-1042"}}, timeout=60)
print(r.status_code)How to decide
Write down the requirement first. If it is 'audio generated inside a named region', exclude services that do not offer that choice. If it is 'low latency for a live call', exclude job APIs. If it is 'a finished voiceover that I can verify', the shape of Sume's job, result and receipt is what you need, and region is not what separates the options.
A short pre-purchase checklist
Ask four questions of any speech API before you commit. Can I name the region? Is the audio delivered as a stream or as a finished file? What is the maximum input per request? What proves which text was spoken? For Sume the answers are: no region choice, a finished file via a job, 20,000 characters per transcript, and a SHA-256 receipt on source-bound jobs. For a live-call product the second answer disqualifies it; for a narrated-video product the fourth answer is the one that matters.
Sources
Related posts
More in Comparisons
- Draft first: Omni 1.1 Flash 360p or Seedance 2.5 at 480p on Sume
Google says Omni 1.1 Flash's 360p draft mode is up to 60% faster at a third of the cost. What that proves, and where Seedance 2.5 480p fits on Sume.
- Omni extension reads 10 s of context; a Sume chain passes a frame
What Google's scene extension carries into the next 10 seconds, and what Sume carriers hand over instead: a last frame, 3-second clips or image refs.
- Omni scene extension to 40 s: Gemini app plans versus Sume per-second
Google offers Omni scene extension to 40 s in Flow and the Gemini app for Plus, Pro and Ultra. On Sume you pay per second instead; here is what carries over.
- OpenAI tts-1-hd at $30 per million characters vs Cartesia Sonic
tts-1-hd is $30 per million characters and tts-1 is $15. Cartesia Sonic is $38 at list on Scale overage, and Sume's Sonic-3.6 route is $47.50 per million.
Written by Sume