OpenAI TTS: 100+ languages but English-optimized voices; test yours
OpenAI's guide lists 100+ languages yet says its voices are optimized for English. Run a 1-cent test in your language on Sume's TTS router before choosing.

OpenAI's text-to-speech guide says the models support 100+ languages, and in the same passage says its voices are currently optimized for English. Read both halves: the first is a coverage claim, the second is a quality caveat. If your audience is not English-speaking, run a short listening test before you build on any large language count. On Sume, the same test is one router call per language and costs 1 cent each.
What the OpenAI page says, and what it leaves out
The guide lists three models (gpt-4o-mini-tts, tts-1 and tts-1-hd), 13 built-in voices, an instructions parameter, and output formats mp3, opus, aac, flac, wav and pcm. It also requires you to tell end users that the voice is AI-generated. The page I read does not rank languages by quality and does not state an input length limit. So a language being 'supported' tells you it will produce speech, not how native that speech sounds.
| Item | What the page says | What it does not say |
|---|---|---|
| Languages | 100+ | Per-language quality |
| Voices | 13 built-in; optimized for English | Native non-English voices |
| Models | gpt-4o-mini-tts, tts-1, tts-1-hd | Per-model language lists |
| Disclosure | Tell users the voice is AI-generated | A format for the notice |
The test, in your language
Write one sentence of about 150 characters in the target language with a number, a place name and a question. Generate it with a voice whose stored language matches. On Sume that is POST /v1/tts-router/generate with model and language set; if the voice's stored language differs from the request, you get a 409 tts_voice_language_mismatch before any job or charge, which protects you from paying for an accidental English accent. Then listen once, and have a native speaker listen once.
import os, requests
r = requests.post("https://api.sume.com/v1/tts-router/generate",
headers={"x-api-key": os.environ["SUME_API_KEY"]},
json={"model": "sonic-3.6", "language": "de",
"transcript": "Ihre Bestellung 4821 kommt am 12. März an. Passt das für Sie?",
"voice": {"id": os.environ["SUME_VOICE_ID"]},
"mode": "sync", "wait_timeout_seconds": 30}, timeout=60)
print(r.status_code, r.json()["data"]["status_url"])Where Sume stands
Sume's router seeds Cartesia Sonic only, and the Sonic 3.6 page lists 44 languages. That is fewer than OpenAI's 100+, so the right comparison is not the count but whether your language is on the Sonic list and sounds right. If it is not on the list, the router is not the answer today, and you should say so rather than force it.
On cost, a 150-character test line is 150 x 0.00475 = 0.7125 cents and bills as 1 cent. A ten-language audition is 10 cents, which is less than the time it takes to argue about a language count.
Disclosure travels with the audio
OpenAI asks for disclosure, and most platforms do too. Keep a record per clip. Sume stores a metadata object on the job and does not send it to the provider, so a field like ai_voice: true is a safe place to keep that record next to the job id.
Scoring the listening test
Give the native listener a one-line form: numbers correct, name correct, stress natural, accent acceptable. Four ticks is a pass. Keep the audio and the form with the job id, so the decision can be revisited when the vendor ships a new model. A pass in March says little about October.
If two voices pass, pick the cheaper one. On Sume the price is per character, so voices from the same router row cost the same, and the choice becomes purely about sound.
Sources
Related posts
More in Comparisons
- A 'Pareto frontier' speech-to-text claim: plot your own cost and delay
Microsoft says MAI-Transcribe-2-Streaming sits on the Pareto frontier. A frontier needs two axes and your data. Plot Sume STT cost against measured job time.
- File-size limits as average bitrate for a 60-second clip
A 60-second clip can average 3.3 Mbps on free ArtStation, 6.7 on a Linktree background, 25 on Reels and 60 on Truth Social. Full table and script.
- Longest clip per platform and how many Sume trim jobs it takes
Truth Social 15 minutes, Reels API 15, TikTok API 10, Kick and Shorts 3, ArtStation 1. How many trim jobs a 30-minute source needs for each, and the cost.
- Predis Core: 325 images or 33 videos vs per-image pricing on Sume
Predis Core lists 1,300 credits, about 325 images or 33 videos. Divide the plan price and compare to Sume's per-image estimate of $0.2225.
Written by Sume