TTS locale vs language field: getting an en-GB accent
Cartesia's locale field picks a regional accent such as en-GB on Sonic 3.6. Sume's TTS request has a language field only, so pick the accent via the voice.

Cartesia separates language (a base code like en) from locale (a regional code like en-GB) and ties accents to each voice. Sume's TTS request exposes a language field only, so it has no locale or accent field: pick a voice that already speaks the accent you want.
Cartesia facts are from its Multilingual voices guide; Sume facts are from the OpenAPI document behind the API reference. Both were read 2026-09-30.
What does Cartesia mean by locale?
Per the guide, most requests need only language or locale: the voice speaks that language in one of the accents it supports. locale needs sonic-3.6 or newer; on earlier models you use language with a base code like en. A request may set one or the other, never both.
Each voice reports an accents list, and each accent belongs to a locale. If you ask for a locale the voice has no accent for, the voice keeps its original accent. The guide's example is a voice whose only English accent is General American: a request for en-GB falls back to American English rather than failing.
What does Sume accept instead?
The schema describes language as the language the voice speaks the transcript in (BCP-47 / ISO-639, for example ko, ja, en), 2 to 16 characters long. Omitting it defaults to English at the provider. The request schema does not allow extra properties, so locale and accent are not part of the contract.
| Control | Cartesia (Sonic 3.6) | Sume TTS |
|---|---|---|
language | Base code such as en | BCP-47 / ISO-639, 2-16 chars |
locale | Regional code such as en-GB | Not a field |
accent | Per-voice accent id | Not a field |
| Default | Not covered here | English at the provider |
Can I send en-GB as the language?
The length limit allows it, since en-GB is five characters, and the description names BCP-47. The docs do not say what Sume does with a regional subtag, though, so do not assume it selects a British accent. Send en, listen, and if the accent is wrong, change the voice.
const response = await fetch("https://api.sume.com/v1/tts-1.0/generate", {
method: "POST",
headers: {
Authorization: "Bearer " + process.env.SUME_API_KEY,
"Content-Type": "application/json",
"Idempotency-Key": "tts-en-001",
},
body: JSON.stringify({
transcript: "Thanks for holding. Your account is up to date.",
voice: { id: process.env.SUME_VOICE_ID },
language: "en",
}),
});
console.log(await response.json());Can I still use sonic-3.6 on Sume?
Yes, as an explicit id. TTS 1.0 has no engine picker, but POST /v1/tts-router/generate requires a model, and its catalog enum includes sonic-3.6, sonic-3.5 and sonic-3. That choice sets the model, not a locale. The Sonic 3.6 languages post covers the language side in detail.
Sources
Related posts
More in Developers
- TTS reads 1999-2000 wrong: write ranges and fractions as words
Cartesia does not normalize ranges like 1999-2000 or fractions like 2/3. Sume sends your transcript literally, so write them as words before you submit.
- TTS voice-language mismatch warning: confirm, then retry
Sume's TTS warns when a voice's language differs from `language`. No job or charge exists yet; after the user agrees, retry with `confirm_language_mismatch`.
- Twitch 2K clips: trim a 1440p clip without setting output
Video trim's output width and height are limited to 256-2160, so a 2560-wide 1440p clip cannot be conformed. Omit output and the source size is kept.
- Vercel 413 FUNCTION_PAYLOAD_TOO_LARGE with Sume video
Vercel Functions cap request and response bodies at 4.5 MB. Do not proxy a generated video through one; hand the client the media.sume.com URL instead.
Written by Sume