HeyGen's 22 Spanish and 17 Arabic variants vs Sume's language field

HeyGen lists regional variants such as 22 Spanish and 17 Arabic. Sume's TTS takes one BCP-47 language string and does not publish a per-region variant list.

5 min readSume
All posts

What does HeyGen's language page say about variants?

HeyGen's help article on video translation languages lists roughly 130 languages, and for several it lists regional variants: Arabic with 17, English with 14 and Spanish with 22. It says you can translate videos using audio-only translation or lip sync, and notes the list keeps expanding. The page does not say which languages support lip sync versus audio-only, so check that in the product before you promise a market a lip-synced version.

If you are comparing providers on regional accuracy (Mexican versus Castilian Spanish, Gulf versus Levantine Arabic), the number of variants is a proxy for how much the vendor has tuned voices to place. It is not proof of quality; listen to samples in the market you care about.

How does Sume handle language and region?

On Sume the language you speak is a string. The OpenAPI description for TTS says language is a BCP-47 or ISO-639 code between 2 and 16 characters, such as ko, ja or en; it is the language the voice speaks the transcript in, and you should set it for every non-English transcript because an omitted value defaults to English at the provider. Sume infers ko or ja from Hangul-only or kana-only text as a fallback.

We found no list of regional variants in the Sume docs, so a field accepting es-MX is not the same as a promise that the voice will sound Mexican. The accent comes from the voice you pick. Our locale versus language post looks at the en-GB case, and the Sonic 3.6 languages post lists what the TTS engine covers.

What happens if the voice and the language disagree?

Sume warns rather than guessing. If the voice's language does not match the language you asked for, the request returns a mismatch warning, and you retry with confirm_language_mismatch set once you are sure. That prevents a silent English-accented Spanish read, which is the usual failure when someone changes language but keeps an English voice.

A safe routine for each market: pick a voice that matches the target language, set language to the code, run a 10-second sample, listen with a native speaker, then run the full script. The sample costs a fraction of the full job and catches wrong stress, wrong numbers and awkward names early.

What about subtitles in those variants?

Burned-in captions on Sume use the text you provide, so a variant is whatever spelling you typed. The caption language field is only a speech-to-text hint and never changes the style or font. Latin display faces handle Spanish accents; the docs list Hangul faces for Korean and no other scripts. For Arabic, which writes right to left and needs different glyphs, we found no documented support in the caption renderer, so do not plan an Arabic burned-in batch without a test render.

Variants and language fields, read 2026-10-02
QuestionHeyGen help articleSume docs
Language countAbout 130 listedNo single list; set per step
Spanish variants22 listedOne language string; accent follows the voice
Arabic variants17 listedSame; burned-in Arabic not documented
Lip sync optionListed alongside audio-onlyOnly for stills plus audio, not existing footage
Mismatch protectionNot describedWarning plus confirm_language_mismatch

Which should you choose for regional markets?

Before choosing, write down what each market actually needs. A Spanish-language ad for Mexico, Spain and Argentina can share a script in many cases, but currency, slang and a product name sometimes cannot. List the lines that must change per region, translate only those, and reuse the rest. That keeps the number of voice generations, and the review time, lower than a full re-translation per variant.

If the main requirement is a vendor that publishes regional variants and offers lip sync on existing video, HeyGen's page is the stronger match on paper. If you want an API where you control the transcript, pick the voice and render captions in separate jobs, Sume fits, with the caveat that regional placement comes from voice selection and a native-speaker listen, not from a variant menu. In both cases keep a human reviewer for each market before anything is published.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume