Bulgarian, Croatian, Slovak text to speech API: one request each
Cartesia's Sonic 3.6 list includes bg, hr and sk. Run three scripts through Sume TTS 1.0 with a language map. Three 2,000-character jobs quote 30 cents.

A buyer who types "Bulgarian text to speech API" usually has three or four Balkan and Central European markets in one project. Cartesia's Sonic 3.6 page (read 2026-10-05) lists Bulgarian, Croatian and Slovak in its 44 languages. Sume TTS 1.0 runs Sonic 3.6, so you reach all three by changing one field, language, to bg, hr or sk.
Why one request per language
A request carries one language value. Mixing a Bulgarian line and a Slovak line in one transcript means one of them is read with the wrong rules. Keep one script per language, and join the outputs afterwards if you need a single file. The stored guide on mixed-language scripts shows the join.
Code
run is the same submit-and-poll helper used across these posts. The loop sends each script with its own code and an idempotency key per language, so a retry cannot bill twice.
import os, time, requests
B = "https://api.sume.com/v1"
H = {"Authorization": "Bearer " + os.environ["SUME_API_KEY"]}
def run(path, body, key=None):
h = {**H, **({"Idempotency-Key": key} if key else {})}
r = requests.post(B + path, json={**body, "mode": "async"}, headers=h)
r.raise_for_status()
job = r.json()["data"]["job"]["id"]
while not requests.get(f"{B}/jobs/{job}/status", headers=H).json()["data"]["terminal"]:
time.sleep(3)
res = requests.get(f"{B}/jobs/{job}/result", headers=H)
res.raise_for_status()
return res.json()["data"]["result"]
SCRIPTS = {l: open(f"{l}.txt", encoding="utf-8").read() for l in ("bg", "hr", "sk")}
VOICES = {l: os.environ["VOICE_" + l.upper()] for l in SCRIPTS}
for lang, text in SCRIPTS.items():
out = run("/tts-1.0/generate", {"transcript": text, "language": lang,
"voice": {"id": VOICES[lang]}}, key=f"promo-oct-{lang}")
print(lang, out["audio_url"])The cost arithmetic
The price is $0.0475 per 1,000 characters. The quote for each request rounds up to a whole cent. Three scripts of 2,000 characters each cost 2 x $0.0475 = $0.095 apiece, so each quotes at 10 cents and the batch at 30 cents. The same 6,000 characters in one request would be $0.285, rounded to 29 cents. The difference is small, and it is the price of keeping each language correct.
Failure modes to expect
- Voice and language disagree:
409 tts_voice_language_mismatch. Audition the voice, then retry withconfirm_language_mismatch: trueif you accept it. - Forgot
language: the text is read as English, and you pay for it. Sume infers only Korean and Japanese from script. - Script over 20,000 characters: split it at a sentence end first.
Pick voices per language
Keep a small dict of voice ids for each language and record which one won your listening test. The result echoes voice and language, which makes later rebuilds consistent.
Sources
Related posts
More in Developers
- Bulk ad queue says completed: read counts.failed before you ship
A Sume bulk queue is completed when every item is terminal, not when every ad worked. A Python poller that backs off, then lists failed and canceled items.
- Bulk queue replay returns 202 and the old queue: derive the key
A repeated Idempotency-Key on a Sume bulk create returns 202 and the old queue, not 200. Derive the key from the batch: safe retries, deliberate reruns.
- Bulk run 400 with details.index: validate your ad variants first
One bad item makes the whole Sume bulk create return 400 and starts nothing. A Python pre-check for concurrency 1 to 16, 1 to 100 items, and the four item keys.
- Bulk text to speech from a CSV: one job per row, safe retries
Voice 500 CSV rows with Sume TTS 1.0. One idempotency key per row id, retry-after on 429, and the queue-capacity numbers per plan decide how fast it goes.
Written by Sume