OpenAI TTS: 100+ languages but English-optimized voices; test yours

OpenAI's guide lists 100+ languages yet says its voices are optimized for English. Run a 1-cent test in your language on Sume's TTS router before choosing.

5 min readSume
All posts

OpenAI's text-to-speech guide says the models support 100+ languages, and in the same passage says its voices are currently optimized for English. Read both halves: the first is a coverage claim, the second is a quality caveat. If your audience is not English-speaking, run a short listening test before you build on any large language count. On Sume, the same test is one router call per language and costs 1 cent each.

What the OpenAI page says, and what it leaves out

The guide lists three models (gpt-4o-mini-tts, tts-1 and tts-1-hd), 13 built-in voices, an instructions parameter, and output formats mp3, opus, aac, flac, wav and pcm. It also requires you to tell end users that the voice is AI-generated. The page I read does not rank languages by quality and does not state an input length limit. So a language being 'supported' tells you it will produce speech, not how native that speech sounds.

OpenAI TTS guide facts that bear on non-English use, read 2026-10-05
ItemWhat the page saysWhat it does not say
Languages100+Per-language quality
Voices13 built-in; optimized for EnglishNative non-English voices
Modelsgpt-4o-mini-tts, tts-1, tts-1-hdPer-model language lists
DisclosureTell users the voice is AI-generatedA format for the notice

The test, in your language

Write one sentence of about 150 characters in the target language with a number, a place name and a question. Generate it with a voice whose stored language matches. On Sume that is POST /v1/tts-router/generate with model and language set; if the voice's stored language differs from the request, you get a 409 tts_voice_language_mismatch before any job or charge, which protects you from paying for an accidental English accent. Then listen once, and have a native speaker listen once.

import os, requests
r = requests.post("https://api.sume.com/v1/tts-router/generate",
    headers={"x-api-key": os.environ["SUME_API_KEY"]},
    json={"model": "sonic-3.6", "language": "de",
          "transcript": "Ihre Bestellung 4821 kommt am 12. März an. Passt das für Sie?",
          "voice": {"id": os.environ["SUME_VOICE_ID"]},
          "mode": "sync", "wait_timeout_seconds": 30}, timeout=60)
print(r.status_code, r.json()["data"]["status_url"])

Where Sume stands

Sume's router seeds Cartesia Sonic only, and the Sonic 3.6 page lists 44 languages. That is fewer than OpenAI's 100+, so the right comparison is not the count but whether your language is on the Sonic list and sounds right. If it is not on the list, the router is not the answer today, and you should say so rather than force it.

On cost, a 150-character test line is 150 x 0.00475 = 0.7125 cents and bills as 1 cent. A ten-language audition is 10 cents, which is less than the time it takes to argue about a language count.

Disclosure travels with the audio

OpenAI asks for disclosure, and most platforms do too. Keep a record per clip. Sume stores a metadata object on the job and does not send it to the provider, so a field like ai_voice: true is a safe place to keep that record next to the job id.

Scoring the listening test

Give the native listener a one-line form: numbers correct, name correct, stress natural, accent acceptable. Four ticks is a pass. Keep the audio and the form with the job id, so the decision can be revisited when the vendor ships a new model. A pass in March says little about October.

If two voices pass, pick the cheaper one. On Sume the price is per character, so voices from the same router row cost the same, and the choice becomes purely about sound.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume