Czech, Hungarian, Romanian, Polish TTS: what MAI lists, what Sume tags
MAI-Voice-2.1 lists cs, hu, ro and pl. Sume tags only pl. A one-cent test plan and the Romanian s-comma versus s-cedilla encoding trap that breaks reads.

Of the four Central and Eastern European languages on Microsoft's MAI-Voice-2.1 list (Czech, Hungarian, Romanian, Polish), Sume's voice library tags only Polish. For the other three, run a one-cent audition on Sume, and fix your text encoding first.
What Microsoft lists
Microsoft's model page lists Czech, Hungarian, Romanian and Polish within its 23 languages, and the OpenRouter Flash page carries cs-CZ, hu-HU, ro-RO and pl-PL.
| Locale code | Listed on | Note |
|---|---|---|
| cs-CZ | MAI-Voice-2.1 and Flash | Czech; not a Sume tag |
| hu-HU | MAI-Voice-2.1 and Flash | Hungarian; not a Sume tag |
| ro-RO | MAI-Voice-2.1 and Flash | Romanian; not a Sume tag |
| pl-PL | MAI-Voice-2.1 and Flash | Polish; also Sume tag pl |
What Sume ships
Sume's voice library tags pl and none of the other three. Polish therefore gets the full guard; Czech, Hungarian and Romanian are audition territory.
Sume's TTS Router (POST /v1/tts-router/generate) takes a required model (sonic-3.6, sonic-3.5, sonic-3, sonic-latest or sonic-preview), one of transcript or transcript_source, a voice selector and a language string of 2 to 16 characters. It is character-metered at $0.0475 per 1,000 characters, rounded up to whole cents per job, with a 20,000-character cap per request.
A one-cent audition
Write one line of about 200 characters that contains your hardest words: a price, a date, a brand name and a number. Save it as line.txt, then submit it and read the job. A line of 210 characters or fewer is the minimum one-cent job. Set language to the code you will use in production; if the chosen voice is tagged for a different language you get the 409 guard rather than a charge.
JOB=$(curl -sS -X POST https://api.sume.com/v1/tts-router/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: audition-ro-001" \
-d "$(jq -n --rawfile t line.txt '{model:"sonic-3.6",transcript:$t,avatar_handle:"product_host",language:"ro"}')" \
| jq -r '.data.request_id')
curl -sS https://api.sume.com/v1/jobs/$JOB/status -H "Authorization: Bearer $SUME_API_KEY"
curl -sS https://api.sume.com/v1/jobs/$JOB/result -H "Authorization: Bearer $SUME_API_KEY"What to listen for
- Hungarian double-acute letters (the o and u with two accents) and Czech r-caron are easy to lose in a spreadsheet export. Check the bytes, not the screen.
- Romanian has two code points for the same-looking letter: s with comma below (U+0219) and s with cedilla (U+015F). Old fonts and exports produce the cedilla form. Normalise to the comma form.
- Polish numerals and declension: read a price and a quantity in the audition line.
- Brand names in Latin script inside the sentence.
- Mixed-language product names, which these languages inflect.
Find the wrong Romanian letters
Run your script through a replace before you send it. The snippet shows the two code points and swaps the cedilla forms to the comma forms.
text = "\u015fi \u0163ara"
print([hex(ord(c)) for c in "\u0219\u015f"])
fixed = text.replace("\u015f", "\u0219").replace("\u0163", "\u021b")
print(fixed)What it costs
Same arithmetic as every other language: 450 characters is 3 cents, 1,200 characters is 6 cents.
| Test | Characters per line | Jobs | Sume cost |
|---|---|---|---|
| One line | 210 or fewer | 1 | $0.01 |
| 6 lines (one per hard case) | 210 or fewer | 6 | $0.06 |
| One 1,000-character script | 1,000 | 1 | $0.05 |
Where this stops
Sume's TTS is asynchronous: you submit, poll GET /v1/jobs/:id/status, then read /result. It does not stream audio and it has no SSML field, so pauses and emphasis come from your punctuation and the generation_config controls (speed 0.6 to 1.5, volume 0.5 to 2, a short emotion guide). Microsoft's Flash model is the one built for live conversation; for a rendered ad or a narrated video, a finished file is the thing you need.
Sources
Related posts
More in Models
- Danish, Finnish, Norwegian TTS: on MAI's list, Sume tags only Swedish
MAI-Voice-2.1 lists da-DK, fi-FI, nb-NO and sv-SE. Sume's voice library tags only sv. How to test a Nordic ad read for 1 cent and keep Bokmal text consistent.
- Dutch text to speech API for ads: MAI-Voice-2.1 nl-NL vs Sume nl
Dutch is on both lists: MAI-Voice-2.1 (nl-NL) and Sume's voice tags (nl). Test numbers, loanwords and the je/u register for 1 cent, then price a campaign.
- Filipino and Cantonese: fil and yue codes in MAI-Transcribe-2 vs Sume
MAI-Transcribe-2-Streaming lists fil and yue and treats unknown codes as unset. Sume takes a 2-16 character language_code hint; test a sample before relying.
- FLUX 3 Image open weights are 'coming weeks': API now or wait?
Press says FLUX 3 Image open weights arrive in the coming weeks; BFL lists commercial weights by licence. Call an API now if you need results this month.
Written by Sume