Latin American Spanish text to speech: es or es-MX on Sume?

Sume's TTS language is a free string and its voice library tags voices with plain es. What that means for Mexican, Argentine or Spain Spanish, and how to test.

5 min readSume
All posts

For Latin American Spanish on Sume, send language: "es", pick a voice you have listened to, and write the script in the vocabulary of your market. Sume's reference does not document regional language codes such as es-MX, so a regional accent comes from the voice and the words, not from a documented setting.

That comes from the API reference for TTS 1.0, read 2026-10-03, plus how Sume's voice-language check behaves. Where the docs are silent, this post says so and gives a test instead of a guess.

What does the language field accept?

The language field is a string of 2 to 16 characters, described as BCP-47 or ISO-639, with ko, ja and en as examples. Sixteen characters leaves room for es-MX, so the request will not be rejected for its length. The reference does not say that a regional subtag changes the voice, and it publishes no list of supported regions.

The mismatch check works on the primary part only. When you use a voice from your Voices library, Sume lowercases the code and takes what comes before a hyphen or underscore, so es-MX is compared as es. A voice tagged es will not trigger a warning for es-MX, and the library's own tags are plain two-letter codes. There is no es-MX voice tag to match against.

So how do you get a Mexican or Argentine sound?

A voice made for a region, from the Voices area of the app, is the only lever Sume documents for accent. Voice cloning is an app feature rather than a field on the TTS request.

  • Choose the voice by ear. Generate the same short line with two or three es voices and keep the one that sounds like your audience.
  • Write the script for the market. Vocabulary and forms of address carry the region as much as the voice does.
  • Send es as the baseline. If you also want to try es-MX, run the same line both ways and compare; the docs promise no difference.
  • Keep one voice per market in your own config, so a Mexican campaign never silently picks the Spain voice.

Script choices that change by market

Spanish wording to settle per market before TTS, read 2026-10-03
ChoiceSpain-leaningLatin America-leaning
Addressing a groupvosotrosustedes
Mobile phonemóvilcelular (many markets)
Computerordenadorcomputadora (many markets)
Carcochecarro or auto, by country
Currency and numberseuro, comma decimalspeso or dollar, check the decimal mark per country

Does a dub need the same voice in every Spanish market?

Not necessarily, and often not. A Mexican audience and an audience in Spain can hear the same voice differently, so a small test with each market's reviewers is cheaper than a recut. Because Sume's TTS charges per character, producing two takes of a 600-character script is a matter of cents, so test both before you commit to one voice for the campaign.

If the video is a dub rather than a new voiceover, keep the timing in mind as well. Spanish usually needs more words than English for the same idea, which can push a line past the slot it replaces. Shorten the Spanish script, then generate again; the per-sentence slices Sume can return with timestamps.words and sentence segmentation show you which line ran over.

What Sume does not do here

Sume does not translate the script in the TTS call, does not publish a region list for Spanish, and does not take a reference recording on the TTS request. Write or buy the localized script, choose the voice by listening, and use the confirmation flag only on purpose. If a regional accent is a hard requirement for a brand voice, create the voice in the app first and use its id.

Record the decision once per market, and include the voice id, the language string you sent and the date. Voices and defaults can change, and a note of what was sent is the fastest way to explain why a re-run sounds different from the version approved last month.

What to ask your reviewer to check

Have a native speaker from the target market listen to one line with a number, a date, and your brand name. Those are where a mismatched voice shows. If the brand name is read wrongly, TTS 1.0 accepts a pronunciation_dict_id, which is the documented way to fix a word without rewriting the sentence.

The table's wording column is general Spanish usage, not a Sume feature; it varies by country, so treat it as a prompt for your reviewer, not a rule. For the engine-side comparison with another vendor's regional tags, see Gemini TTS regional dialects. The wrong-accent case from a missing field is in the English-accent post.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume