Syllable-controlled translation for voiceover that must fit
Translated speech often runs longer than the original. Index-Homura targets length-controlled translation. How to budget words against clip time before TTS.

When translated speech has to fit the same clip length, translate with a length budget and check the result against the time you have, before you generate any audio. Bilibili's Index-Translate family lists a part called Index-Homura for syllable-controlled translation, built for that dubbing problem (README read 2026-10-04). It is an open component of the family and we have not benchmarked it.
Whatever model you use, the underlying issue is the same. Many language pairs expand: the same sentence takes more syllables, so the voiceover overruns the picture.
Why does translated speech run long?
Speaking rate is roughly a syllables-per-second figure that changes by language and speaker. If the target language needs more syllables for the same meaning, you either speed the voice up until it sounds rushed, or you cut the line. A length-aware translation does the cutting upstream, in the text, which keeps the voice natural.
| Step | Without length control | With a length budget |
|---|---|---|
| Translate | Free-length text | Text aimed at a syllable target |
| Check | Often skipped | Compare estimated duration with the slot |
| Voice | Speed up or overrun | Natural pace, fits the slot |
| Picture | Cut or freeze to match | Unchanged |
How do you budget a line?
The same discipline applies to on-screen text: a reading-speed check flags cues that are too long for their time.
- Take the slot in seconds from the cue's start and end.
- Multiply by a comfortable speaking rate for your target language to get a character or syllable cap. Measure that rate on a sample of your own voice, not a guess.
- Ask the translator for a line at or under the cap, and re-ask for any line that goes over.
- Only then synthesise the voice.
How does this relate to a Sume avatar script?
Sume's avatar videos accept a script when Sume estimates the video duration at 4 to 60 seconds inclusive. A translated script that grows past 60 seconds is rejected, so a length budget protects you from a failed request as well as from a rushed voice. Trim the translation, or split the script into scenes, before you submit.
For standalone text to speech, the same budget applies; a hosted comparison is in the open TTS post.
What if the line is still too long?
Do not stretch the voice past a natural pace to rescue a long line. Trim the line instead: drop filler words, prefer the shorter synonym, or merge it with the previous cue's idea so the total meaning survives in fewer syllables. If meaning cannot survive, give that cue a longer slot by moving the cut in your edit rather than speeding the voice.
Keep a log of lines you had to trim. Patterns show up quickly, and they tell you whether your target language needs a tighter budget everywhere.
Sources
Related posts
More in Developers
- Free Index-Translate API for client subtitles: what to check
Index-Translate's free endpoint is labelled an online demo with no stated rate limit. A checklist before client transcripts go in, and the self-host option.
- Keep brand names intact in translated, burned-in captions
Index-Translate supports glossary instructions, but vendor notes say compliance is not guaranteed. Check every term in code before Sume burns the cues in.
- Translate all subtitle cues in one JSON request, then validate
Index-Translate lists JSON format preservation. Send a clip's cues as one JSON array, then validate count, timings and length before Sume burns them.
- Translate a long transcript in 60-second chunks for captions
Index-Translate's FP8 card sets a 4096-token limit and Sume caption jobs cover 60 seconds. Split cues into windows, translate each, and shift times to zero.
Written by Sume