Nova 2 Sonic runs in 4 AWS regions; can you pick one on Sume TTS?

Amazon lists Nova 2 Sonic GA in N. Virginia, Oregon, Tokyo and Stockholm. Sume's TTS has one base URL and no region field; here is what that means for you.

5 min readSume
All posts

No: on Sume you cannot pick a region for text-to-speech. The request schema has no region field, and every call goes to one base URL, https://api.sume.com. Amazon's Nova 2 release notes list general availability for Nova 2 Sonic in four AWS regions: US East (N. Virginia), US West (Oregon), Tokyo and Stockholm. If your rule says audio must be generated in a named region, that is a fact to settle before you choose an API, and a job-based service with one endpoint will not satisfy it.

What the Amazon page lists

The same release notes describe two refreshes. A March 2026 refresh added Polly-compatible voices, cut p50 latency by 150 ms and improved turn-taking on 8 kHz telephony audio. A May 2026 refresh reported 88% fewer hallucinations, 52% less speaker drift and 28% fewer critical errors, each on an internal data set, and was deployed in place between May 21 and 28 with no API change. Regions matter because a real-time voice agent wants the model close to the caller.

Amazon Nova 2 Sonic regions and refreshes versus the Sume TTS surface, read 2026-10-05
QuestionAmazon release notesSume TTS 1.0 / router
Where does it run?US East (N. Virginia), US West (Oregon), Tokyo, StockholmOne base URL, https://api.sume.com; no region field
Telephony audioMarch 2026: better 8 kHz turn-takingoutput_format sample_rate can be 8000
Model selectionModel refreshed in placeRouter model id chosen per request
DeliveryLive sessionsAsync job; sync wait up to 30 s; webhook

When a region matters, and when it does not

Region matters for two reasons: latency for live speech, and data residency rules. Prepared audio, such as a narrated video, is not latency-bound, because nobody waits for the first byte. Data residency is a policy question, and the honest answer for Sume is that a request is handled through its API without a region choice, so if your policy names a region, check with your own compliance owner before you send text containing personal data.

Sume does give you ways to limit what is sent. The metadata object you attach to a TTS job is stored with the job and is not sent to the provider, so keep customer ids there instead of in the transcript.

import os, requests
r = requests.post("https://api.sume.com/v1/tts-router/generate",
    headers={"x-api-key": os.environ["SUME_API_KEY"]},
    json={"model": "sonic-3.6",
          "transcript": "Your appointment is at 4 PM on Tuesday.",
          "voice": {"id": os.environ["SUME_VOICE_ID"]},
          "output_format": {"container": "wav", "sample_rate": 8000,
                            "encoding": "pcm_mulaw"},
          "metadata": {"customer_ref": "c-1042"}}, timeout=60)
print(r.status_code)

How to decide

Write down the requirement first. If it is 'audio generated inside a named region', exclude services that do not offer that choice. If it is 'low latency for a live call', exclude job APIs. If it is 'a finished voiceover that I can verify', the shape of Sume's job, result and receipt is what you need, and region is not what separates the options.

A short pre-purchase checklist

Ask four questions of any speech API before you commit. Can I name the region? Is the audio delivered as a stream or as a finished file? What is the maximum input per request? What proves which text was spoken? For Sume the answers are: no region choice, a finished file via a job, 20,000 characters per transcript, and a SHA-256 receipt on source-bound jobs. For a live-call product the second answer disqualifies it; for a narrated-video product the fourth answer is the one that matters.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume