Tagalog text to speech API: Sume tl voices vs Eleven v4 Filipino
Sume's Voices library tags voices with tl for Tagalog; Eleven v4 lists Filipino among 90+ languages. What to send, and how to check the voice speaks it.

For Tagalog text to speech, Sume's Voices library has a Tagalog language tag, tl, and ElevenLabs' Eleven v4 lists Filipino in its language set. Both can voice a Tagalog script; the difference is how each tells the model which language to speak. On Sume you pick a voice tagged tl and send language: "tl" with the transcript.
Sume's side comes from the tts_create payload described in MCP tools and gates and the Voices language list in the Sume codebase. ElevenLabs' side comes from its Eleven v4 docs page and launch post, both read on 2026-10-03.
What each side lists
The ElevenLabs docs page lists Filipino in its language list and says some languages are handled more fluently than others. Sume's Voices library accepts 16 language tags when you create a voice, and tl is one of them. A Sume voice carries one language tag, so you pick or create a voice for Tagalog rather than switching a single voice between languages.
| Item | Sume | ElevenLabs Eleven v4 |
|---|---|---|
| Tagalog or Filipino listed | Yes, tag tl in the Voices library | Yes, Filipino in the language list |
| Where you say the language | language field on the TTS request, plus the voice's own tag | Model chooses from the text, per the docs page |
| Number of languages | 16 voice language tags | 90+ on the launch post |
| Wrong-language guard | Warning and confirm_language_mismatch before any charge | Not stated on the pages read |
The request that works
Send the Tagalog transcript, a voice id from your Voices library whose language is tl, and language: "tl". The language field matters: the tool schema says the provider assumes English when it is omitted. If the voice's tag and the request language disagree, Sume returns a warning with would_submit: false and creates no job until you resend with confirm_language_mismatch: true.
The transcript can be up to 20,000 characters per request, and spaces and punctuation count toward the price.
Check before you publish
Neither vendor's page promises equal quality across every language, and ElevenLabs says so directly. Generate a 200-character test, listen for stress and borrowed English vowels, and only then run the full script. Tagalog scripts that mix English words are the usual trouble spot; try the exact mixed sentence you plan to ship.
Sources
Related posts
More in Comparisons
- Veo 3.1 Lite vs Wan 3.0: price per second at 480p, 720p and 1080p
Google lists Veo 3.1 Lite at $0.05 (720p), $0.08 (1080p) per second; fal lists Wan 3.0 at $0.05 (480p), $0.10 (720p), $0.20 (1080p).
- Veo 3.1 tiers vs Seedance 2.5: price per second of video
Google lists Veo 3.1 at $0.40 (Standard), $0.10-$0.12 (Fast) and $0.05-$0.08 (Lite) per second; Seedance 2.5 on fal is about $0.46 at 720p. Sume x 1.25: $0.58.
- Wan 3.0 or MiniMax H3 Max: which to pin for reference-to-video
Both take image, video and audio references on Sume. Wan runs 2 to 30 seconds with 5 reference videos; H3 Max runs 5 to 15 seconds with always-on stereo audio.
- Wan 3.0 or Seedance 2.5: which is cheaper for a 10-second 720p clip?
Wan 3.0 is cheaper: a 10-second 720p clip is $1.00 at fal list ($1.25 on Sume) against about $4.62 ($5.78 on Sume) for Seedance 2.5, roughly 4.6x.
Written by Sume