Bulgarian, Croatian, Slovak text to speech API: one request each

Cartesia's Sonic 3.6 list includes bg, hr and sk. Run three scripts through Sume TTS 1.0 with a language map. Three 2,000-character jobs quote 30 cents.

4 min readSume
All posts

A buyer who types "Bulgarian text to speech API" usually has three or four Balkan and Central European markets in one project. Cartesia's Sonic 3.6 page (read 2026-10-05) lists Bulgarian, Croatian and Slovak in its 44 languages. Sume TTS 1.0 runs Sonic 3.6, so you reach all three by changing one field, language, to bg, hr or sk.

Why one request per language

A request carries one language value. Mixing a Bulgarian line and a Slovak line in one transcript means one of them is read with the wrong rules. Keep one script per language, and join the outputs afterwards if you need a single file. The stored guide on mixed-language scripts shows the join.

Code

run is the same submit-and-poll helper used across these posts. The loop sends each script with its own code and an idempotency key per language, so a retry cannot bill twice.

import os, time, requests
B = "https://api.sume.com/v1"
H = {"Authorization": "Bearer " + os.environ["SUME_API_KEY"]}

def run(path, body, key=None):
    h = {**H, **({"Idempotency-Key": key} if key else {})}
    r = requests.post(B + path, json={**body, "mode": "async"}, headers=h)
    r.raise_for_status()
    job = r.json()["data"]["job"]["id"]
    while not requests.get(f"{B}/jobs/{job}/status", headers=H).json()["data"]["terminal"]:
        time.sleep(3)
    res = requests.get(f"{B}/jobs/{job}/result", headers=H)
    res.raise_for_status()
    return res.json()["data"]["result"]

SCRIPTS = {l: open(f"{l}.txt", encoding="utf-8").read() for l in ("bg", "hr", "sk")}
VOICES = {l: os.environ["VOICE_" + l.upper()] for l in SCRIPTS}
for lang, text in SCRIPTS.items():
    out = run("/tts-1.0/generate", {"transcript": text, "language": lang,
              "voice": {"id": VOICES[lang]}}, key=f"promo-oct-{lang}")
    print(lang, out["audio_url"])

The cost arithmetic

The price is $0.0475 per 1,000 characters. The quote for each request rounds up to a whole cent. Three scripts of 2,000 characters each cost 2 x $0.0475 = $0.095 apiece, so each quotes at 10 cents and the batch at 30 cents. The same 6,000 characters in one request would be $0.285, rounded to 29 cents. The difference is small, and it is the price of keeping each language correct.

Failure modes to expect

  • Voice and language disagree: 409 tts_voice_language_mismatch. Audition the voice, then retry with confirm_language_mismatch: true if you accept it.
  • Forgot language: the text is read as English, and you pay for it. Sume infers only Korean and Japanese from script.
  • Script over 20,000 characters: split it at a sentence end first.

Pick voices per language

Keep a small dict of voice ids for each language and record which one won your listening test. The result echoes voice and language, which makes later rebuilds consistent.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume