Transcribe a MAI-Voice-2.1 clip with Sume STT: Python audio_url run

Host a MAI-Voice-2.1 or Flash clip at a public HTTPS URL and Sume STT returns text, word times and sentences for 1 cent a minute. A 25-line Python run.

5 min readSume
All posts

Put the clip at a public HTTPS URL and POST it to /v1/stt-1.0/transcribe. Sume returns the text, a words[] array with start and end seconds, and sentence segments[] when you ask. Because MAI-Voice-2.1-Flash makes at most 45 seconds of audio per call (Microsoft AI, read 2026-10-05), one clip is well inside the 10-minute cap of an STT job.

This is a useful check of a synthetic take: you hear what the model said, in text, with times you can use for captions.

What the request takes

The request schema asks for a public HTTPS audio_url, preferably on the Sume media host. language_code is an optional hint, and omitting it means auto-detect. duration_seconds runs from 1 to 600 and improves the usage reservation; omit it and Sume reserves one minute. Word timings are always returned, so there is no flag for them. Send segmentation: {"mode": "sentence"} to also get sentences. diarize and tag_audio_events are fixed on the server, and sending them is rejected.

A Python run

Every submit needs an Idempotency-Key. Poll the job status with backoff until it is terminal, then read the result, as Jobs and results describes. The script reads the key from SUME_API_KEY, so it stops with a clear error when the variable is missing.

Sume's docs call the submit response's request_id the job id, so the script reads that and falls back to id.

import os, time, uuid, requests

BASE = "https://api.sume.com"
KEY = os.environ["SUME_API_KEY"]
H = {"Authorization": f"Bearer {KEY}"}

def transcribe(audio_url, seconds):
    body = {"audio_url": audio_url, "duration_seconds": seconds,
            "segmentation": {"mode": "sentence"}}
    r = requests.post(f"{BASE}/v1/stt-1.0/transcribe", json=body,
                      headers={**H, "Idempotency-Key": str(uuid.uuid4())})
    r.raise_for_status()
    j = r.json()
    job = j.get("request_id") or j["id"]
    delay = 2
    while True:
        s = requests.get(f"{BASE}/v1/jobs/{job}/status", headers=H).json()
        if s.get("status") in ("completed", "failed", "canceled"):
            break
        time.sleep(delay)
        delay = min(delay * 2, 15)
    res = requests.get(f"{BASE}/v1/jobs/{job}/result", headers=H).json()
    return s.get("status"), res

if __name__ == "__main__":
    print(transcribe("https://example.com/take.wav", 40))

Before you run it

The URL in the last line is a placeholder. Use your own HTTPS file. The audio field expects a public HTTPS URL, so a file on a private network or a localhost server will not work.

What it costs

STT is billed per audio minute: Sume lists $0.01 a minute, which is $0.60 an hour. A 40-second clip reserves about 0.7 cents. Microsoft's streaming transcription model is listed at $0.54 an hour through the end of 2026 and works over a live connection; a finished 40-second file does not need that.

Do not treat the transcript as a quality score. It tells you what a recognizer heard. Compare it with the script yourself, as in the round-trip check linked below.

Reading the result

Read the result for text, language_code, words[] with word, start and end, and segments[] with index, text, start, end and duration_seconds. Segments are gapless: each one ends where the next begins, and they are time ranges over your file, not sliced audio. Words are capped at 20,000 entries, which a 600-second job stays well under, and a capped result says so with words_truncated and words_total.

For a 45-second clip that is a few hundred timed tokens at most, which is easy to diff against the script you sent to the voice. A simple check is to join the word fields, lower-case them, and compare with the lower-cased script. A missing or doubled word is the usual sign that the voice skipped or repeated text, and the timings tell you where to listen.

Run it on a handful of takes before you transcribe a whole catalog, and keep the audio you test at the same sample rate you will ship, because a clip resampled for the test can hide a problem the real file has.

Keep the check cheap and routine. At one cent a minute a re-check of every take is a rounding error next to the render that follows.

When it goes wrong

Two failures are common. A URL that needs a login returns a fetch error, because the route expects a public HTTPS file. A clip with no speech returns empty or near-empty text, which is itself a signal that the synthesis went wrong. A transport error while polling is not a job outcome: read the status again and never submit the paid request a second time without the same idempotency key.

Sume's errors page lists 429 rate_limited with a retry hint and queue_full when the workspace has no queue room. Back off on both and keep the key.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume