Japanese and Tagalog TTS: not on MAI-Voice-2.1, tagged on Sume
MAI-Voice-2.1's 23 languages exclude Japanese and Tagalog; Sume's voice library tags ja and tl. How to set language and audition for 1 cent.

If you need Japanese or Tagalog speech from the same week's launch news, MAI-Voice-2.1 and MAI-Voice-2.1-Flash will not help: neither language is on Microsoft's list of 23. Sume's voice library tags both ja and tl, and the TTS Router takes them through the language field.
What Microsoft lists
Microsoft's MAI-Voice-2.1 page names 23 languages. The OpenRouter page for the Flash model lists 28 locale codes. Japanese and Filipino/Tagalog are on neither list, so for these two languages the launch gives you nothing to compare against.
| Locale code | Listed on | Note |
|---|---|---|
| ja-JP | Not listed | Absent from Microsoft's 23 languages and from the 28 locale codes on OpenRouter |
| fil-PH / tl-PH | Not listed | Absent from both lists |
| ja | Sume voice tag | Sume voice library |
| tl | Sume voice tag | fil and tl are treated as the same language |
What Sume ships
Sume's voice library tags ja and tl among its 16 languages. Sume also guesses Japanese or Korean from a transcript that is only kana or only Hangul, but only as a fallback: kanji-only text has no such hint, so always send language: "ja" for Japanese. For Tagalog, fil and tl are accepted as the same language when the guard compares them.
Sume's TTS Router (POST /v1/tts-router/generate) takes a required model (sonic-3.6, sonic-3.5, sonic-3, sonic-latest or sonic-preview), one of transcript or transcript_source, a voice selector and a language string of 2 to 16 characters. It is character-metered at $0.0475 per 1,000 characters, rounded up to whole cents per job, with a 20,000-character cap per request.
A one-cent audition
Write one line of about 200 characters that contains your hardest words: a price, a date, a brand name and a number. Save it as line.txt, then submit it and read the job. A line of 210 characters or fewer is the minimum one-cent job. Set language to the code you will use in production; if the chosen voice is tagged for a different language you get the 409 guard rather than a charge.
JOB=$(curl -sS -X POST https://api.sume.com/v1/tts-router/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: audition-ja-001" \
-d "$(jq -n --rawfile t line.txt '{model:"sonic-3.6",transcript:$t,avatar_handle:"product_host",language:"ja"}')" \
| jq -r '.data.request_id')
curl -sS https://api.sume.com/v1/jobs/$JOB/status -H "Authorization: Bearer $SUME_API_KEY"
curl -sS https://api.sume.com/v1/jobs/$JOB/result -H "Authorization: Bearer $SUME_API_KEY"What to listen for
- Numbers and units: ask for a price such as 12,800 yen and a date, and check each is read the way a Japanese listener expects.
- Kanji with several readings: if a place or person name is read wrongly, the
pronunciation_dict_idfield accepts a pronunciation dictionary id; otherwise rewrite the name in kana. - Mixed script: Latin brand names inside Japanese text are the most common failure; keep them in the audition line.
- Tagalog and English code-switching: Taglish copy should be audition-tested as written, not translated to pure Tagalog first.
- Speed: Japanese copy often needs
generation_config.speedbelow 1.0 for a calm read; the range is 0.6 to 1.5.
What it costs
A Japanese ad script of 300 characters is a 2-cent job; a 1,200-character narration is 6 cents. Microsoft's per-million rates (the $22 and $15 figures) cannot be applied here because the language is not offered.
| Test | Characters per line | Jobs | Sume cost |
|---|---|---|---|
| One line | 210 or fewer | 1 | $0.01 |
| 8 lines (one per hard case) | 210 or fewer | 8 | $0.08 |
| One 1,000-character script | 1,000 | 1 | $0.05 |
Where this stops
Sume's TTS is asynchronous: you submit, poll GET /v1/jobs/:id/status, then read /result. It does not stream audio and it has no SSML field, so pauses and emphasis come from your punctuation and the generation_config controls (speed 0.6 to 1.5, volume 0.5 to 2, a short emotion guide). Microsoft's Flash model is the one built for live conversation; for a rendered ad or a narrated video, a finished file is the thing you need.
Sources
Related posts
More in Models
- Kling 3.0 native 4K at 60 fps: what Sume's kling-3 lets you order
Kling 3.0 is marketed with native 4K. The fal page lists up to 1080p, and Sume's kling-3 accepts 720p or 1080p, 4 to 15 s. Here is what to plan around.
- kling-3 has no reference_video_urls; motion_video_url is another route
Sume's kling-3 id rejects reference_*_urls. A clip that should drive a character's motion goes to the separate Kling motion control route as motion_video_url.
- 'kling-3 supports text-to-video or start/end frames only': the fix
kling-3 on Sume has no reference lists. The 400 is fixed by sending text only, or a first frame and optional end frame, or by switching to a reference model.
- Kling 4.0 Auto aspect ratio vs Sume's explicit aspect_ratio field
Kling 4.0 lists an Auto format next to 16:9, 1:1, 9:16 and 21:9. Sume's video ids take explicit ratios; see which ids list which values and how to read them.
Written by Sume