Budget a translated script per sentence before you dub it

Turn each source sentence's seconds into a character budget before translating, so the dubbed read fits the slot. A method with Sume sentence segments.

5 min readSume
All posts

To keep a dub inside its slot, give your translator a character budget for each sentence: the sentence's seconds from the source segments multiplied by your measured characters per second. Sume's sentence segments supply the seconds, and the TTS word timings let you measure the rate.

Why per sentence

A translated script that is two seconds too long overall may be fine on 18 sentences and badly wrong on two. Compare totals and you will learn the dub overruns but not where. Compare per sentence and the translator can tighten exactly the sentences that fail.

Sume's STT and video inspect both return gapless sentence segments when you ask for segmentation: { mode: "sentence" }, where each segment ends exactly where the next starts. Video inspect only returns them with transcribe on, and STT bills $0.01 per audio minute.

Measure the rate on your voice

Do not use a number from a blog post. Render a 300-character test paragraph in the target language with timestamps.words: true, divide the characters by the end time of the last word, and use the result. The 15 characters a second used elsewhere in these posts is an assumption for English; other languages will differ, and voices differ again.

Then the budget is arithmetic. A source sentence that lasts 6.0 seconds, read in a language that your test shows at 14 characters a second, gets 84 characters.

A per-sentence character budget at 14 characters a second (worked example; Sume sentence segments, read 2026-10-05)
SentenceSource secondsBudget in charactersTranslation lengthVerdict
14.05652Fits
26.08497Tighten by 13
33.54947Fits
45.07074Tighten by 4

What to do with an overrun

Retakes are cheap. A 97-character line is 1 cent, the minimum charge, so redo it until it fits rather than argue about it.

  • Edit the text first. A shorter sentence is always cheaper than a faster voice.
  • Raise speed only a little. The range is 0.6 to 1.5, and a read near the top sounds rushed.
  • Spend the pause. A source with a long gap after a sentence can absorb a longer line, so check the next segment's start.

A simple checker

The check is a few lines in any language, and the point is to run it before the translator is paid for a second pass.

import json, sys

rate = float(sys.argv[1])  # measured characters per second
rows = json.load(open("sentences.json"))
bad = 0
for i, row in enumerate(rows, 1):
    budget = int((row["end"] - row["start"]) * rate)
    n = len(row["translation"])
    if n > budget:
        bad += 1
        print(f"sentence {i}: {n} > {budget}, cut {n - budget}")
print("overruns:", bad)

A note on languages

Microsoft's MAI-Voice-2.1 page lists 23 languages (read 2026-10-05), and a language list says whether a voice exists, not how long a given sentence runs. Budgeting is how you find out. Run it per language, because a German sentence and a Spanish sentence of the same meaning rarely have the same length.

Where to keep the numbers

Keep the sentence table as a file in the project, with the source seconds, the budget, the translation and the verdict in one row. When the client changes a line, you edit one row and rerun the check, and the audio retake for that row is 1 cent. The cost of getting this right is mostly discipline, not money.

Share the budget column with the translator, not just the source script. A translator who sees 84 next to a sentence will write to 84, and the fixes in the table above never happen.

Limits of the method

A character budget is an estimate. Numbers, acronyms and names read differently from ordinary words, and a line full of them can run long or short against the rate. Treat a sentence within about ten percent of its budget as a candidate for a listen rather than a certain pass or fail, and trust the word timings from the rendered audio as the final check.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume