Benchmark TTS latency yourself: submit to finished file in Python

Vendors quote 45 ms or 150 ms; your pipeline waits for a file. A Python harness that times Sume TTS jobs from submit to a finished state, with p50 and p95.

6 min readSume
All posts

The latency number that matters is the one you measure on your own scripts, from your own network, to the point you can use the output. Microsoft lists MAI-Voice-2.1-Flash at about 45 ms of model inference and 150 ms end to end for 45 seconds of audio, and MAI-Voice-2.1 at about 550 ms (read 2026-10-03). Those are vendor figures for a streaming-style service. A Sume TTS job is a different shape: you submit, the job queues, runs and completes, and you read a finished file.

This harness measures that shape: time to accepted, time to terminal state, and the spread across many runs. You can point the same timing code at any TTS service.

Two different latencies

The Sume docs state that for paid generation queued is a normal accepted state, and workspace concurrency limits apply when workers move jobs into processing. So a burst of jobs will show queue time, and that is a property of your plan and load, not the model.

Vendor numbers from Microsoft AI's MAI-Voice-2.1 page and launch post, read 2026-10-03; job states from Sume's Jobs and results page.
MeasureWhat it tells youWho quotes it
Time to first audioHow soon playback can startStreaming vendors (about 45 ms inference, 150 ms end to end for Flash)
Time to complete fileWhen a render step can proceedYou, with a harness
Queue timeWait before work startsJob APIs: queued is a normal state
Tail latency (p95)What a batch really waits forOnly you can measure this

The harness

It submits one TTS Router job per run, polls /v1/jobs/:id/status with exponential backoff, and records seconds until completed, failed or canceled. Set SUME_API_KEY and a voice id in SUME_VOICE_ID first. It uses a unique idempotency key per run and never resubmits a request that is still running, which follows the docs' advice not to resubmit paid work because a local process timed out.

import json
import os
import statistics
import time
import urllib.request
import uuid

BASE = "https://api.sume.com"
KEY = os.environ["SUME_API_KEY"]
VOICE = os.environ["SUME_VOICE_ID"]

def call(method, path, body=None, headers=None):
    data = json.dumps(body).encode() if body is not None else None
    h = {"Authorization": "Bearer " + KEY, "Content-Type": "application/json"}
    h.update(headers or {})
    req = urllib.request.Request(BASE + path, data=data, headers=h, method=method)
    with urllib.request.urlopen(req, timeout=60) as r:
        return json.load(r)

def one_run(text):
    t0 = time.monotonic()
    job = call("POST", "/v1/tts-router/generate",
               {"model": "sonic-3.6", "transcript": text, "voice": {"id": VOICE}, "language": "en"},
               {"Idempotency-Key": "bench-" + uuid.uuid4().hex})
    job_id = job.get("id") or job.get("request_id")
    accepted = time.monotonic() - t0
    delay = 0.5
    while True:
        status = call("GET", f"/v1/jobs/{job_id}/status")
        state = status.get("status") or status.get("job", {}).get("status")
        if state in ("completed", "failed", "canceled"):
            return accepted, time.monotonic() - t0, state
        time.sleep(delay)
        delay = min(delay * 1.5, 5.0)

if __name__ == "__main__":
    runs = [one_run("Meet the new travel mug. It fits every cup holder.") for _ in range(10)]
    totals = sorted(r[1] for r in runs)
    print("states:", sorted({r[2] for r in runs}))
    print("p50 %.2fs  p95 %.2fs" % (statistics.median(totals), totals[int(0.95 * (len(totals) - 1))]))

Make the comparison fair

Compare like with like. Use the same text, the same length and the same language for every engine, run each at least ten times, and alternate engines rather than finishing one before starting the next, so a slow minute on your network hits both. Log the time of day. Record the voice used, because voice choice can change audio length and, for a job API, finish time.

Keep the cost beside the time. A table with p50, p95 and dollars per 1,000 characters lets you see that a result 0.6 seconds slower is also 30 percent cheaper, or that it is not. For Sume's router the rate is list price times 1.25, which is $47.50 per 1M characters; Microsoft lists $15 and $22 per 1M for the two MAI voice models. Those are the three numbers your table should carry.

Finally, check the output, not just the clock. A fast job that returns a clipped ending is slow once you count the retake. Listen to the last second of every run, and mark any that end early.

Reading the numbers

Run it at the hour you will really use it, with scripts as long as your real ones, and include a few ten-job bursts if you batch. Report the p50 and p95 together; the median hides the tail that decides how long a nightly batch takes. If p95 is far above p50, look at queueing before blaming the model.

Then decide by use. If a person is waiting for the first syllable, a job API is the wrong tool and a streaming service is right. If your next step is a render, a caption or a join, you need the file, and a few seconds of job time matters less than getting the voice, the text and the cost right. Sume's TTS docs describe no streaming; the Router's non-goals list streaming TTS as out of scope.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume