Measure your voice's characters per second, then budget video slots

Guessing words per minute breaks a 3-second hook. Voice one Sume take, divide characters by duration_seconds, and size each video slot from your own number.

4 min readSume
All posts

To fit a line into a 3-second slot, voice one real take and divide its character count by duration_seconds from the job result. That number is characters per second for that voice, language and speed, and it replaces any rule of thumb. Sume returns duration_seconds on every completed TTS job, and one 500-character calibration take costs $0.02375 (API reference).

The script

Use text that looks like your real copy, not a pangram. A mix of short and long sentences averages out. Spaces and punctuation count as characters for billing, so count them the same way here.

import os, time, requests
API = "https://api.sume.com"
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}

def run(path, body, key):
    r = requests.post(API + path, json=body, timeout=60,
                      headers={**H, "Idempotency-Key": key})
    r.raise_for_status()
    job = r.json()["request_id"]
    while True:
        s = requests.get(f"{API}/v1/jobs/{job}/status", headers=H, timeout=30).json()
        if s.get("terminal"):
            break
        time.sleep(s.get("next_poll_after_seconds") or 3)
    res = requests.get(f"{API}/v1/jobs/{job}/result", headers=H, timeout=30)
    res.raise_for_status()
    return res.json()

sample = open("calibration.txt").read().strip()
res = run("/v1/tts-1.0/generate", {
    "transcript": sample,
    "avatar_handle": os.environ["VOICE_HANDLE"],
    "output_format": {"container": "wav", "encoding": "pcm_s16le",
                      "sample_rate": 44100},
}, "calibrate-v1")
cps = len(sample) / res["duration_seconds"]
print(f"{cps:.2f} characters per second")
for slot in (3, 5, 8):
    print(slot, "s slot fits about", int(slot * cps), "characters")

Use it as a budget, not a promise

The figure holds for one voice, one language and one speed. Re-measure when you change any of them. generation_config.speed runs from 0.6 to 1.5, and a change there moves the number, so measure again after each change. Treat the slot sizes as a first draft, then voice the real line and check its duration_seconds against the slot.

When the take is too long

You have three levers, in this order. Cut words, which is free. Then raise generation_config.speed a little and listen for strain. Last, move the next shot's start on the timeline. The render takes video[].start as authoritative, and it requires the first slot to start at 0 (timeline docs).

Why measure per voice

Every new voice or model changes the pace. The new releases make that a routine event: Microsoft's Flash tier claims 55 percent faster inference, which says nothing about speaking rate (Microsoft AI, read 2026-10-04). Speed of generation and speed of speech are different numbers. Keep a small table of characters per second by voice and language beside your preset, and update it when you change either one.

A worked example

Suppose a calibration take of 600 characters returns a duration_seconds of 40. That is 15 characters per second for that voice, so a 3-second hook holds about 45 characters and an 8-second product line about 120. These numbers are an example, not a benchmark: your voice, language and speed will give a different figure, which is the reason to measure. Keep the calibration text in your repo and rerun it for $0.0285 whenever a voice or setting changes.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume