Budget a translated script per sentence before you dub it
Turn each source sentence's seconds into a character budget before translating, so the dubbed read fits the slot. A method with Sume sentence segments.

To keep a dub inside its slot, give your translator a character budget for each sentence: the sentence's seconds from the source segments multiplied by your measured characters per second. Sume's sentence segments supply the seconds, and the TTS word timings let you measure the rate.
Why per sentence
A translated script that is two seconds too long overall may be fine on 18 sentences and badly wrong on two. Compare totals and you will learn the dub overruns but not where. Compare per sentence and the translator can tighten exactly the sentences that fail.
Sume's STT and video inspect both return gapless sentence segments when you ask for segmentation: { mode: "sentence" }, where each segment ends exactly where the next starts. Video inspect only returns them with transcribe on, and STT bills $0.01 per audio minute.
Measure the rate on your voice
Do not use a number from a blog post. Render a 300-character test paragraph in the target language with timestamps.words: true, divide the characters by the end time of the last word, and use the result. The 15 characters a second used elsewhere in these posts is an assumption for English; other languages will differ, and voices differ again.
Then the budget is arithmetic. A source sentence that lasts 6.0 seconds, read in a language that your test shows at 14 characters a second, gets 84 characters.
| Sentence | Source seconds | Budget in characters | Translation length | Verdict |
|---|---|---|---|---|
| 1 | 4.0 | 56 | 52 | Fits |
| 2 | 6.0 | 84 | 97 | Tighten by 13 |
| 3 | 3.5 | 49 | 47 | Fits |
| 4 | 5.0 | 70 | 74 | Tighten by 4 |
What to do with an overrun
Retakes are cheap. A 97-character line is 1 cent, the minimum charge, so redo it until it fits rather than argue about it.
- Edit the text first. A shorter sentence is always cheaper than a faster voice.
- Raise
speedonly a little. The range is 0.6 to 1.5, and a read near the top sounds rushed. - Spend the pause. A source with a long gap after a sentence can absorb a longer line, so check the next segment's start.
A simple checker
The check is a few lines in any language, and the point is to run it before the translator is paid for a second pass.
import json, sys
rate = float(sys.argv[1]) # measured characters per second
rows = json.load(open("sentences.json"))
bad = 0
for i, row in enumerate(rows, 1):
budget = int((row["end"] - row["start"]) * rate)
n = len(row["translation"])
if n > budget:
bad += 1
print(f"sentence {i}: {n} > {budget}, cut {n - budget}")
print("overruns:", bad)A note on languages
Microsoft's MAI-Voice-2.1 page lists 23 languages (read 2026-10-05), and a language list says whether a voice exists, not how long a given sentence runs. Budgeting is how you find out. Run it per language, because a German sentence and a Spanish sentence of the same meaning rarely have the same length.
Where to keep the numbers
Keep the sentence table as a file in the project, with the source seconds, the budget, the translation and the verdict in one row. When the client changes a line, you edit one row and rerun the check, and the audio retake for that row is 1 cent. The cost of getting this right is mostly discipline, not money.
Share the budget column with the translator, not just the source script. A translator who sees 84 next to a sentence will write to 84, and the fixes in the table above never happen.
Limits of the method
A character budget is an estimate. Numbers, acronyms and names read differently from ordinary words, and a line full of them can run long or short against the rate. Treat a sentence within about ten percent of its budget as a candidate for a listen rather than a certain pass or fail, and trust the word timings from the rendered audio as the final check.
Sources
Related posts
More in Use cases
- Reel speed 0.5x to 3x: burn captions after the final speed change
Instagram lets creators change Reel speed from 0.5x to 3x. Burned-in captions speed up with the clip, so a 1.5 s line is 0.5 s at 3x. Burn them last.
- Call center roleplay training: customer persona clips with avatars
Record angry or confused customer prompts for agent practice with Sume avatars: one avatar per video, 10-second scripts, cost of a 24-clip scenario bank.
- Can a 30-second AI clip be a YouTube Short? Yes, up to 3 minutes
YouTube Help says a Short can be up to 3 minutes, so a 30-second AI clip fits. How to pick the model length and cut it to the file you upload, with Sume docs.
- Cancellation save video with an AI avatar: one 30-second clip, one ask
A cancellation-save clip should be 30 seconds and make one offer. At Plus it costs $7.35 to render; send it only to accounts where a human would also reach out.
Written by Sume