Cartesia accent field: multilingual voices only; Sume uses voice id

Cartesia's accent field is for multilingual voices only and works independent of locale. Sume's TTS request has no accent field: pick a voice id and language.

4 min readSume
All posts

Cartesia's changelog describes a new optional accent field "for multilingual voices only" that sets how the voice sounds while speaking, independent of locale. Sume's TTS 1.0 request documents no accent field; the sound of the result comes from the voice id you send and the language you set.

Where does each choice live in the two APIs?

Cartesia splits the choice across fields; Sume folds it into the voice and the language. Read 2026-10-01.

Where the sound of a voice is chosen, from the Cartesia changelog and Sume API reference, read 2026-10-01.
ChoiceCartesiaSume TTS 1.0
How the voice soundsaccent, multilingual voices onlyThe voice id
Which voices offer itAccents field on Get VoiceThe voice you pick
Level, pace, feelingNot covered heregeneration_config volume, speed, emotion

How do I choose a voice on Sume?

Send a voice id, or an avatar reference that Sume resolves to its voice. The voice id must be a TTS voice UUID or a Voices library id, "not a voice name from another TTS ecosystem". Anything else is rejected with 400 public_reason=invalid_voice_id before a job is queued or credits are reserved. See the API reference.

Where do I set language and delivery?

Set language for non-English transcripts. generation_config is the optional volume, speed and emotion block. For how language is handled, see Sonic 3.6 languages and the language field.

What if I need a specific accent today?

Audition voices and pick the id that sounds right, then pin it. Do not send a Cartesia voice name; Sume will reject it. For where cloning stands, read voice cloning API.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume