Syllable-controlled translation for voiceover that must fit

Translated speech often runs longer than the original. Index-Homura targets length-controlled translation. How to budget words against clip time before TTS.

5 min readSume
All posts

When translated speech has to fit the same clip length, translate with a length budget and check the result against the time you have, before you generate any audio. Bilibili's Index-Translate family lists a part called Index-Homura for syllable-controlled translation, built for that dubbing problem (README read 2026-10-04). It is an open component of the family and we have not benchmarked it.

Whatever model you use, the underlying issue is the same. Many language pairs expand: the same sentence takes more syllables, so the voiceover overruns the picture.

Why does translated speech run long?

Speaking rate is roughly a syllables-per-second figure that changes by language and speaker. If the target language needs more syllables for the same meaning, you either speed the voice up until it sounds rushed, or you cut the line. A length-aware translation does the cutting upstream, in the text, which keeps the voice natural.

Where length control sits in a dubbing chain, read 2026-10-04
StepWithout length controlWith a length budget
TranslateFree-length textText aimed at a syllable target
CheckOften skippedCompare estimated duration with the slot
VoiceSpeed up or overrunNatural pace, fits the slot
PictureCut or freeze to matchUnchanged

How do you budget a line?

The same discipline applies to on-screen text: a reading-speed check flags cues that are too long for their time.

  • Take the slot in seconds from the cue's start and end.
  • Multiply by a comfortable speaking rate for your target language to get a character or syllable cap. Measure that rate on a sample of your own voice, not a guess.
  • Ask the translator for a line at or under the cap, and re-ask for any line that goes over.
  • Only then synthesise the voice.

How does this relate to a Sume avatar script?

Sume's avatar videos accept a script when Sume estimates the video duration at 4 to 60 seconds inclusive. A translated script that grows past 60 seconds is rejected, so a length budget protects you from a failed request as well as from a rushed voice. Trim the translation, or split the script into scenes, before you submit.

For standalone text to speech, the same budget applies; a hosted comparison is in the open TTS post.

What if the line is still too long?

Do not stretch the voice past a natural pace to rescue a long line. Trim the line instead: drop filler words, prefer the shorter synonym, or merge it with the previous cue's idea so the total meaning survives in fewer syllables. If meaning cannot survive, give that cue a longer slot by moving the cut in your edit rather than speeding the voice.

Keep a log of lines you had to trim. Patterns show up quickly, and they tell you whether your target language needs a tighter budget everywhere.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume