Hebrew and Georgian text to speech API: he and ka on Sume TTS

Cartesia lists Hebrew (he) and Georgian (ka) on Sonic 3.6, 3.5 and 3. What to send to Sume TTS, why the library has no voice for them, and what a script costs.

4 min readSume
All posts

Hebrew and Georgian are both spoken by Sume TTS when you send language: "he" or language: "ka" with a voice that speaks them. Cartesia's model pages list he and ka on Sonic 3.6, Sonic 3.5 and Sonic 3, and Sume's TTS 1.0 route uses the current stable Sonic, so there is nothing to opt into. Both use non-Latin scripts, which makes the language field essential: left out, the provider default is English.

The catch is the voice. Neither language is in Sume's voice-library list, so a voice for them is a UUID you obtain, not something Sume's voice tool creates.

Which models list these languages?

The Sonic 2 and Sonic Turbo snapshots list only a handful of languages and neither includes Hebrew or Georgian; both stop working after October 20, 2026 according to Cartesia. The TTS Router's ids (sonic-3.6, sonic-3.5, sonic-3, sonic-latest, sonic-preview) are all Sonic 3 generation or later.

Language codes per Cartesia's model pages (read 2026-10-03); the last column is Sume's voice-library list from its voice code
LanguageCodesonic-3.6-2026-08-27sonic-3.5-2026-05-04sonic-3-2026-01-12In Sume's voice-library language list
HebrewheYesYesYesNo
GeorgiankaYesYesYesNo

How do I handle a voice for a language outside the library?

Sume's voice selector accepts a TTS voice UUID or a voi_ id and rejects other shapes with invalid_voice_id, before a job or charge. A UUID is a Cartesia voice id. Sume's tool guidance says Assets and Voices rows are optional references, not an admission gate, and that library metadata never gates a raw voice id.

So the workflow is: pick or clone a Hebrew or Georgian voice on the provider side, copy its UUID, and send it as voice.id. If Sume knows the voice and its saved language differs, you get the 409 tts_voice_language_mismatch double-check; if it does not know it, you get no check, and listening to the first take is the control.

Is there anything special about right-to-left text?

Sume's contract treats the transcript as a string of 1 to 20,000 characters and says nothing about direction, so send Hebrew the way it is stored and test one take. What you do need to watch is what you do with the audio afterwards: word timings from timestamps.words come back as start and end seconds per word, which is what a caption step consumes. If you plan to burn captions for a right-to-left language, read Sume's caption documentation for the scripts it covers before you commit.

  • Price: $0.0475 per 1,000 characters, 20,000 maximum per request.
  • Audio cap: 1,200 seconds per job, tts_duration_exceeded beyond it.
  • Default output: MP3, 44,100 Hz, 128 kbps.

Request example

Hebrew with a Hebrew voice UUID in VOICE_ID. For Georgian change the text and use ka.

import os
import uuid

import requests

r = requests.post(
    "https://api.sume.com/v1/tts-1.0/generate",
    headers={
        "x-api-key": os.environ["SUME_API_KEY"],
        "Idempotency-Key": str(uuid.uuid4()),
    },
    json={
        "transcript": "ברוכים הבאים לעדכון השבועי שלנו.",
        "language": "he",
        "voice": {"mode": "id", "id": os.environ["VOICE_ID"]},
        "mode": "async",
    },
    timeout=30,
)
r.raise_for_status()
print(r.json())

What should I check on the first take?

Listen for vowel handling in Hebrew, where text is often written without vowel marks, and for loanwords in both languages. A sample of 200 characters costs $0.0095 at Sume's rate, so it is cheap to test a voice before a long script. If a name is read incorrectly, spell it as it should sound in the transcript. Keep one voice id per language so a series sounds consistent.

Sources

Related posts

More in Models

All Models posts

Written by Sume