Spain vs Mexico, Brazil vs Portugal: regional voiceover on MAI, Sume

MAI-Voice-2.1 lists es-ES, es-MX, pt-BR and pt-PT. Sume's guard compares only the primary language. How to pick a regional read and test it for 1 cent.

4 min readSume
All posts

Microsoft lists the regional variants as separate locales (es-ES, es-MX, pt-BR, pt-PT). Sume has one es and one pt tag, and its mismatch guard compares only the primary language, so the region comes from the voice you pick and the copy you write, not from a documented field.

What Microsoft lists

On Microsoft's model page Spanish (Spain), Spanish (Mexico), Portuguese (Brazil) and Portuguese (Portugal) are separate entries. OpenRouter's Flash page lists voices by locale as well.

Regional Spanish and Portuguese locales (read 2026-10-05)
Locale codeListed onNote
es-ESMAI-Voice-2.1 model pageSpanish (Spain)
es-MXMAI-Voice-2.1 model pageSpanish (Mexico)
pt-BRMAI-Voice-2.1 model pagePortuguese (Brazil)
pt-PTMAI-Voice-2.1 model pagePortuguese (Portugal)

What Sume ships

A Sume request may carry a regional tag like es-MX; the guard reads it as es. Nothing in Sume's documentation says the tag changes the accent, so do not assume it does. Choose the voice by listening, and write the copy in the regional register.

Sume's TTS Router (POST /v1/tts-router/generate) takes a required model (sonic-3.6, sonic-3.5, sonic-3, sonic-latest or sonic-preview), one of transcript or transcript_source, a voice selector and a language string of 2 to 16 characters. It is character-metered at $0.0475 per 1,000 characters, rounded up to whole cents per job, with a 20,000-character cap per request.

A one-cent audition

Write one line of about 200 characters that contains your hardest words: a price, a date, a brand name and a number. Save it as line.txt, then submit it and read the job. A line of 210 characters or fewer is the minimum one-cent job. Set language to the code you will use in production; if the chosen voice is tagged for a different language you get the 409 guard rather than a charge.

JOB=$(curl -sS -X POST https://api.sume.com/v1/tts-router/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: audition-esmx-001" \
  -d "$(jq -n --rawfile t line.txt '{model:"sonic-3.6",transcript:$t,avatar_handle:"product_host",language:"es-MX"}')" \
  | jq -r '.data.request_id')
curl -sS https://api.sume.com/v1/jobs/$JOB/status -H "Authorization: Bearer $SUME_API_KEY"
curl -sS https://api.sume.com/v1/jobs/$JOB/result -H "Authorization: Bearer $SUME_API_KEY"

What to listen for

  • Spain: vosotros forms and words such as ordenador; Mexico: ustedes and computadora.
  • Brazil: voce, celular, onibus; Portugal: telemovel, autocarro, and the spelling differences your copy editor knows.
  • Currency and number format: 1.200,50 versus 1,200.50 changes how a price is read.
  • Put one regional word in the audition line so you can hear whether the voice owns it.
  • Keep one regional variant per campaign file; do not mix them in one request.

What it costs

Four regional variants of a 450-character ad read are four separate jobs: 4 x $0.03 = $0.12 on Sume.

Audition cost on Sume at list rate (read 2026-10-05)
TestCharacters per lineJobsSume cost
One line210 or fewer1$0.01
8 lines (one per hard case)210 or fewer8$0.08
One 1,000-character script1,0001$0.05

Where this stops

Sume's TTS is asynchronous: you submit, poll GET /v1/jobs/:id/status, then read /result. It does not stream audio and it has no SSML field, so pauses and emphasis come from your punctuation and the generation_config controls (speed 0.6 to 1.5, volume 0.5 to 2, a short emotion guide). Microsoft's Flash model is the one built for live conversation; for a rendered ad or a narrated video, a finished file is the thing you need.

Sources

Related posts

More in Models

All Models posts

Written by Sume