Czech, Hungarian, Romanian, Greek TTS API: eight codes on Sume

Sonic 3.6 lists cs, sk, hu, ro, bg, hr, el and uk. Set language for Cyrillic and Greek text or the provider reads English. Sume TTS request, codes, and price.

4 min readSume
All posts

Central and Eastern European text works on Sume TTS if you send the right language code: cs Czech, sk Slovak, hu Hungarian, ro Romanian, bg Bulgarian, hr Croatian, el Greek or uk Ukrainian. Cartesia lists all eight on Sonic 3.6, 3.5 and 3. The field matters more here than for Spanish or French, because two of the scripts are not Latin: Bulgarian and Ukrainian are Cyrillic and Greek has its own alphabet, and an omitted language means English at the provider.

Sume's script fallback does not rescue you. It infers only Korean from mostly-Hangul text and Japanese from mostly-kana text, so Cyrillic or Greek text with no language is sent as English.

Which codes are covered and which are in Sume's library?

None of the eight is in Sume's voice-library list (Polish and Russian are, which is the closest the library gets). Synthesis does not need a library row: the voice.id field takes a voice UUID or a voi_ id, and a raw id is never blocked by missing library metadata.

Language codes per Cartesia's model pages (read 2026-10-03); the last column is Sume's voice-library list from its voice code
LanguageCodesonic-3.6-2026-08-27sonic-3.5-2026-05-04sonic-3-2026-01-12In Sume's voice-library language list
CzechcsYesYesYesNo
SlovakskYesYesYesNo
HungarianhuYesYesYesNo
RomanianroYesYesYesNo
BulgarianbgYesYesYesNo
CroatianhrYesYesYesNo
GreekelYesYesYesNo
UkrainianukYesYesYesNo

Is Ukrainian the same as Russian?

No, and the codes keep them apart: uk for Ukrainian, ru for Russian. Sending Ukrainian text with ru is a language mismatch even though both are Cyrillic. If the voice you use is saved in Sume as ru, a request tagged uk returns the 409 tts_voice_language_mismatch double-check before anything is queued; confirm it only if you listened to the voice and accept its accent.

How should I batch a multi-country release?

Cartesia describes one language per request, and Sume's contract says the same: set language to the language the voice speaks the transcript in. A release in eight languages is eight requests, each with its own voice id and its own idempotency key, followed by a join if you want one file. Sume's audio timeline can join up to 20 gapless parts.

Price is language-independent at $0.0475 per 1,000 characters, so the cost difference between languages is only script length. Eight 2,000-character scripts are 16,000 characters, or $0.76 in total.

Cost at Sume's published rate (Sume API OpenAPI document, read 2026-10-03)
ScriptsCharactersCost
1 x 2,0002,000$0.095
4 x 2,0008,000$0.38
8 x 2,00016,000$0.76

Request example

Greek with a Greek-capable voice UUID in VOICE_ID:

import os
import uuid

import requests

r = requests.post(
    "https://api.sume.com/v1/tts-1.0/generate",
    headers={
        "x-api-key": os.environ["SUME_API_KEY"],
        "Idempotency-Key": str(uuid.uuid4()),
    },
    json={
        "transcript": "Καλώς ήρθατε στην εβδομαδιαία ενημέρωσή μας.",
        "language": "el",
        "voice": {"mode": "id", "id": os.environ["VOICE_ID"]},
        "mode": "async",
    },
    timeout=30,
)
r.raise_for_status()
print(r.json())

What should I check on the first take?

Listen for stress and loanwords. Cartesia lists the languages, but it does not promise that every voice reads every domain term well, so a test sample of 200 characters ($0.0095 at Sume's price) is the right first call. If a term is read badly, rewrite it phonetically in the transcript or use a pronunciation dictionary id. Keep one voice per language across a campaign so the series sounds consistent.

Sources

Related posts

More in Models

All Models posts

Written by Sume