Hebrew and Georgian text to speech API: he and ka on Sume TTS
Cartesia lists Hebrew (he) and Georgian (ka) on Sonic 3.6, 3.5 and 3. What to send to Sume TTS, why the library has no voice for them, and what a script costs.

Hebrew and Georgian are both spoken by Sume TTS when you send language: "he" or language: "ka" with a voice that speaks them. Cartesia's model pages list he and ka on Sonic 3.6, Sonic 3.5 and Sonic 3, and Sume's TTS 1.0 route uses the current stable Sonic, so there is nothing to opt into. Both use non-Latin scripts, which makes the language field essential: left out, the provider default is English.
The catch is the voice. Neither language is in Sume's voice-library list, so a voice for them is a UUID you obtain, not something Sume's voice tool creates.
Which models list these languages?
The Sonic 2 and Sonic Turbo snapshots list only a handful of languages and neither includes Hebrew or Georgian; both stop working after October 20, 2026 according to Cartesia. The TTS Router's ids (sonic-3.6, sonic-3.5, sonic-3, sonic-latest, sonic-preview) are all Sonic 3 generation or later.
| Language | Code | sonic-3.6-2026-08-27 | sonic-3.5-2026-05-04 | sonic-3-2026-01-12 | In Sume's voice-library language list |
|---|---|---|---|---|---|
| Hebrew | he | Yes | Yes | Yes | No |
| Georgian | ka | Yes | Yes | Yes | No |
How do I handle a voice for a language outside the library?
Sume's voice selector accepts a TTS voice UUID or a voi_ id and rejects other shapes with invalid_voice_id, before a job or charge. A UUID is a Cartesia voice id. Sume's tool guidance says Assets and Voices rows are optional references, not an admission gate, and that library metadata never gates a raw voice id.
So the workflow is: pick or clone a Hebrew or Georgian voice on the provider side, copy its UUID, and send it as voice.id. If Sume knows the voice and its saved language differs, you get the 409 tts_voice_language_mismatch double-check; if it does not know it, you get no check, and listening to the first take is the control.
Is there anything special about right-to-left text?
Sume's contract treats the transcript as a string of 1 to 20,000 characters and says nothing about direction, so send Hebrew the way it is stored and test one take. What you do need to watch is what you do with the audio afterwards: word timings from timestamps.words come back as start and end seconds per word, which is what a caption step consumes. If you plan to burn captions for a right-to-left language, read Sume's caption documentation for the scripts it covers before you commit.
- Price: $0.0475 per 1,000 characters, 20,000 maximum per request.
- Audio cap: 1,200 seconds per job,
tts_duration_exceededbeyond it. - Default output: MP3, 44,100 Hz, 128 kbps.
Request example
Hebrew with a Hebrew voice UUID in VOICE_ID. For Georgian change the text and use ka.
import os
import uuid
import requests
r = requests.post(
"https://api.sume.com/v1/tts-1.0/generate",
headers={
"x-api-key": os.environ["SUME_API_KEY"],
"Idempotency-Key": str(uuid.uuid4()),
},
json={
"transcript": "ברוכים הבאים לעדכון השבועי שלנו.",
"language": "he",
"voice": {"mode": "id", "id": os.environ["VOICE_ID"]},
"mode": "async",
},
timeout=30,
)
r.raise_for_status()
print(r.json())What should I check on the first take?
Listen for vowel handling in Hebrew, where text is often written without vowel marks, and for loanwords in both languages. A sample of 200 characters costs $0.0095 at Sume's rate, so it is cheap to test a voice before a long script. If a name is read incorrectly, spell it as it should sound in the transcript. Keep one voice id per language so a series sounds consistent.
Sources
Related posts
More in Models
- Higgsfield Soul 2 custom_reference_id is not a Sume parameter
Higgsfield's Soul 2 API takes custom_reference_id for a saved character. Sume's Soul is text-only; for a consistent character use a model that takes references.
- Soul 2 lookbook batch of 4: Higgsfield Soul on Sume
Generate a four-image lookbook set per request with Higgsfield Soul 2: batch_size 1 or 4, 720p or 1080p, seven ratios, text-to-image only.
- Higgsfield Soul on Sume: text to image, 1 or 4 images, 720p or 1080p
Soul is a text-to-image row in Sume's image catalog: seven aspect ratios, no references, 1 or 4 images per call, 720p or 1080p. Limits and price.
- How many Seedance 2.5 clips fill a 3-minute Reel: six at 30 s
Seedance 2.5 on Sume takes 4-30 s per clip, so a 3-minute Reel needs at least six. The clip math, the stitch step and Instagram's 3-minute guidance.
Written by Sume