Japanese and Tagalog TTS: not on MAI-Voice-2.1, tagged on Sume

MAI-Voice-2.1's 23 languages exclude Japanese and Tagalog; Sume's voice library tags ja and tl. How to set language and audition for 1 cent.

4 min readSume
All posts

If you need Japanese or Tagalog speech from the same week's launch news, MAI-Voice-2.1 and MAI-Voice-2.1-Flash will not help: neither language is on Microsoft's list of 23. Sume's voice library tags both ja and tl, and the TTS Router takes them through the language field.

What Microsoft lists

Microsoft's MAI-Voice-2.1 page names 23 languages. The OpenRouter page for the Flash model lists 28 locale codes. Japanese and Filipino/Tagalog are on neither list, so for these two languages the launch gives you nothing to compare against.

Japanese and Tagalog coverage (read 2026-10-05)
Locale codeListed onNote
ja-JPNot listedAbsent from Microsoft's 23 languages and from the 28 locale codes on OpenRouter
fil-PH / tl-PHNot listedAbsent from both lists
jaSume voice tagSume voice library
tlSume voice tagfil and tl are treated as the same language

What Sume ships

Sume's voice library tags ja and tl among its 16 languages. Sume also guesses Japanese or Korean from a transcript that is only kana or only Hangul, but only as a fallback: kanji-only text has no such hint, so always send language: "ja" for Japanese. For Tagalog, fil and tl are accepted as the same language when the guard compares them.

Sume's TTS Router (POST /v1/tts-router/generate) takes a required model (sonic-3.6, sonic-3.5, sonic-3, sonic-latest or sonic-preview), one of transcript or transcript_source, a voice selector and a language string of 2 to 16 characters. It is character-metered at $0.0475 per 1,000 characters, rounded up to whole cents per job, with a 20,000-character cap per request.

A one-cent audition

Write one line of about 200 characters that contains your hardest words: a price, a date, a brand name and a number. Save it as line.txt, then submit it and read the job. A line of 210 characters or fewer is the minimum one-cent job. Set language to the code you will use in production; if the chosen voice is tagged for a different language you get the 409 guard rather than a charge.

JOB=$(curl -sS -X POST https://api.sume.com/v1/tts-router/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: audition-ja-001" \
  -d "$(jq -n --rawfile t line.txt '{model:"sonic-3.6",transcript:$t,avatar_handle:"product_host",language:"ja"}')" \
  | jq -r '.data.request_id')
curl -sS https://api.sume.com/v1/jobs/$JOB/status -H "Authorization: Bearer $SUME_API_KEY"
curl -sS https://api.sume.com/v1/jobs/$JOB/result -H "Authorization: Bearer $SUME_API_KEY"

What to listen for

  • Numbers and units: ask for a price such as 12,800 yen and a date, and check each is read the way a Japanese listener expects.
  • Kanji with several readings: if a place or person name is read wrongly, the pronunciation_dict_id field accepts a pronunciation dictionary id; otherwise rewrite the name in kana.
  • Mixed script: Latin brand names inside Japanese text are the most common failure; keep them in the audition line.
  • Tagalog and English code-switching: Taglish copy should be audition-tested as written, not translated to pure Tagalog first.
  • Speed: Japanese copy often needs generation_config.speed below 1.0 for a calm read; the range is 0.6 to 1.5.

What it costs

A Japanese ad script of 300 characters is a 2-cent job; a 1,200-character narration is 6 cents. Microsoft's per-million rates (the $22 and $15 figures) cannot be applied here because the language is not offered.

Audition cost on Sume at list rate (read 2026-10-05)
TestCharacters per lineJobsSume cost
One line210 or fewer1$0.01
8 lines (one per hard case)210 or fewer8$0.08
One 1,000-character script1,0001$0.05

Where this stops

Sume's TTS is asynchronous: you submit, poll GET /v1/jobs/:id/status, then read /result. It does not stream audio and it has no SSML field, so pauses and emphasis come from your punctuation and the generation_config controls (speed 0.6 to 1.5, volume 0.5 to 2, a short emotion guide). Microsoft's Flash model is the one built for live conversation; for a rendered ad or a narrated video, a finished file is the thing you need.

Sources

Related posts

More in Models

All Models posts

Written by Sume