Text to speech in Japanese: set the language to ja
For Japanese text to speech, send the script in kana and kanji and set the language to ja. Left out, Sume can read a kanji-only line as English.

To convert Japanese text to speech, write the script in Japanese characters (kana and kanji) and set the language to Japanese, ja, explicitly, so it is read as Japanese rather than with English rules. With the Sume API that is one POST /v1/tts-1.0/generate request carrying the Japanese transcript, a voice, and language: "ja". Leave the language out and Sume guesses Japanese only from kana, so a line written in kanji alone, or in romaji, gets the English default.
The request fields come from the TTS schema in the Sume API reference, the OpenAPI document behind the API reference docs, read on 2026-09-28. Checks called current behavior are read from Sume's code, and the price from the code behind API pricing. The rest of the request works as for any language (text to speech API); Korean has its own checks (text to speech in Korean).
How do I send a Japanese text to speech request?
Put the script in transcript exactly as it should be read, in kana and kanji, and add language: "ja". The voice is chosen as for any language, with a ready avatar's avatar_handle or a voice.id you already hold, and the finished job returns a Sume-hosted audio file, MP3 by default.
curl -X POST https://api.sume.com/v1/tts-1.0/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: tts-ja-welcome-001" \
-d '{
"transcript": "新しいマグカップです。本日発売。",
"voice": { "id": "voi_0123456789abcdef0123456789abcdef" },
"language": "ja"
}'What happens if I leave out the language?
Omitted, language defaults to English at the provider; the schema says Sume infers ja only as a fallback, from a kana-only transcript. In current code that fallback counts kana (hiragana and katakana) against Hangul and Latin letters and returns ja when kana outnumber both. Kanji are not counted at all.
| Transcript | `language` sent | Language used |
|---|---|---|
| 新しいマグカップです。 | ja | Japanese, as sent. |
| 新しいマグカップです。 | Omitted | ja, inferred: 9 kana, no Latin letters. |
| 新製品を発売 | Omitted | ja, inferred: 1 kana (を) is enough when nothing else is counted. |
| 本日発売 | Omitted | English: all kanji, so nothing is inferred. |
| Sumeで音声を作る | Omitted | English: 3 kana against 4 Latin letters. |
| konnichiwa | Omitted | English: romaji has no kana. |
Can I send romaji instead of Japanese characters?
The API won't refuse it, but don't count on it. In current code the only script check before a job is for Korean: language: "ko" with no Hangul syllable fails with 400 tts_language_script_mismatch. Nothing checks Japanese text against ja, and the docs don't describe how romaji sent with ja is read. Without ja, romaji has no kana, so it is read as English. Write the script in kana and kanji.
Which voice should read Japanese?
One recorded in Japanese. The API reference publishes no list of Japanese voices; in current code each voice in Sume's Voices library records one of 16 languages, ja among them. When a library voice's language differs from the request's, the submit stops with a 409 tts_voice_language_mismatch double-check before any job or charge, as text to speech pronunciation explains.
- The Japanese catch: with
languageomitted, current code checks a line it can't infer as English, so a Japanese library voice reading a kanji-only line stops at that double-check. Sendingjaavoids it. - In current code, a voice outside the library is not checked, and an avatar's own voice is cloned with its language set to English. Listen to one short Japanese line before you send a long script.
How is Japanese text to speech billed?
Japanese has no rate of its own: text to speech costs $0.0475 per 1,000 characters, plus a 5.5% agent fee by default, counted on transcript characters with spaces and punctuation included. In current code the count is the transcript's length, where a kana such as あ or a kanji such as 東 is one character, and one request takes up to 20,000 of them.
What doesn't Sume's text to speech do for Japanese?
- It publishes no control over the reading of a kanji or its pitch accent.
pronunciation_dict_idis accepted, but no way to create a dictionary is documented; text to speech pronunciation covers what you can change. - Avatar 1.0 talking videos speak English only in current code. For Japanese speech on a face, make the audio here and lip-sync a still image to it, as in lip sync API.
- The Video captions docs describe styles for Latin and Korean wording only, so don't plan on burned-in Japanese captions.
Sources
Related posts
More in Models
- Video to anime AI: how to turn a clip into anime
To turn a video into anime with AI, use a video-to-video edit that names the style and what to keep, or restyle one frame and animate it.
- Wan 3.0 vs Seedance 2.5: 30-second clips, inputs and price
Wan 3.0 and Seedance 2.5 both make clips of up to 30 s with audio, frames and references. Wan starts at 2 s and bills per second, Seedance per token.
- An OpenRouter-compatible video API: sume/auto or a pinned model
Sume's POST /v1/videos follows OpenRouter's video generation API field for field. Let sume/auto pick the model, or pin a catalog id like seedance-2.5.
- Image generation API with reference images: POST /v1/images
Send a prompt plus public HTTPS reference images to Sume's POST /v1/images. Pin a catalog model or send sume/auto; the catalog lists each model's limits.
Written by Sume