Transcribe a MAI-Voice-2.1 clip with Sume STT: Python audio_url run
Host a MAI-Voice-2.1 or Flash clip at a public HTTPS URL and Sume STT returns text, word times and sentences for 1 cent a minute. A 25-line Python run.

Put the clip at a public HTTPS URL and POST it to /v1/stt-1.0/transcribe. Sume returns the text, a words[] array with start and end seconds, and sentence segments[] when you ask. Because MAI-Voice-2.1-Flash makes at most 45 seconds of audio per call (Microsoft AI, read 2026-10-05), one clip is well inside the 10-minute cap of an STT job.
This is a useful check of a synthetic take: you hear what the model said, in text, with times you can use for captions.
What the request takes
The request schema asks for a public HTTPS audio_url, preferably on the Sume media host. language_code is an optional hint, and omitting it means auto-detect. duration_seconds runs from 1 to 600 and improves the usage reservation; omit it and Sume reserves one minute. Word timings are always returned, so there is no flag for them. Send segmentation: {"mode": "sentence"} to also get sentences. diarize and tag_audio_events are fixed on the server, and sending them is rejected.
A Python run
Every submit needs an Idempotency-Key. Poll the job status with backoff until it is terminal, then read the result, as Jobs and results describes. The script reads the key from SUME_API_KEY, so it stops with a clear error when the variable is missing.
Sume's docs call the submit response's request_id the job id, so the script reads that and falls back to id.
import os, time, uuid, requests
BASE = "https://api.sume.com"
KEY = os.environ["SUME_API_KEY"]
H = {"Authorization": f"Bearer {KEY}"}
def transcribe(audio_url, seconds):
body = {"audio_url": audio_url, "duration_seconds": seconds,
"segmentation": {"mode": "sentence"}}
r = requests.post(f"{BASE}/v1/stt-1.0/transcribe", json=body,
headers={**H, "Idempotency-Key": str(uuid.uuid4())})
r.raise_for_status()
j = r.json()
job = j.get("request_id") or j["id"]
delay = 2
while True:
s = requests.get(f"{BASE}/v1/jobs/{job}/status", headers=H).json()
if s.get("status") in ("completed", "failed", "canceled"):
break
time.sleep(delay)
delay = min(delay * 2, 15)
res = requests.get(f"{BASE}/v1/jobs/{job}/result", headers=H).json()
return s.get("status"), res
if __name__ == "__main__":
print(transcribe("https://example.com/take.wav", 40))Before you run it
The URL in the last line is a placeholder. Use your own HTTPS file. The audio field expects a public HTTPS URL, so a file on a private network or a localhost server will not work.
What it costs
STT is billed per audio minute: Sume lists $0.01 a minute, which is $0.60 an hour. A 40-second clip reserves about 0.7 cents. Microsoft's streaming transcription model is listed at $0.54 an hour through the end of 2026 and works over a live connection; a finished 40-second file does not need that.
Do not treat the transcript as a quality score. It tells you what a recognizer heard. Compare it with the script yourself, as in the round-trip check linked below.
Reading the result
Read the result for text, language_code, words[] with word, start and end, and segments[] with index, text, start, end and duration_seconds. Segments are gapless: each one ends where the next begins, and they are time ranges over your file, not sliced audio. Words are capped at 20,000 entries, which a 600-second job stays well under, and a capped result says so with words_truncated and words_total.
For a 45-second clip that is a few hundred timed tokens at most, which is easy to diff against the script you sent to the voice. A simple check is to join the word fields, lower-case them, and compare with the lower-cased script. A missing or doubled word is the usual sign that the voice skipped or repeated text, and the timings tell you where to listen.
Run it on a handful of takes before you transcribe a whole catalog, and keep the audio you test at the same sample rate you will ship, because a clip resampled for the test can hide a problem the real file has.
Keep the check cheap and routine. At one cent a minute a re-check of every take is a rounding error next to the render that follows.
When it goes wrong
Two failures are common. A URL that needs a login returns a fetch error, because the route expects a public HTTPS file. A clip with no speech returns empty or near-empty text, which is itself a signal that the synthesis went wrong. A transport error while polling is not a job outcome: read the status again and never submit the paid request a second time without the same idempotency key.
Sume's errors page lists 429 rate_limited with a retry hint and queue_full when the workspace has no queue room. Back off on both and keep the key.
Sources
Related posts
More in Developers
- Translate a pack into 8 languages in parallel: queue limits by plan
Eight Ideogram 4.5 edits fit the Pro queue (24 accepted jobs) but not Free (6), so 2 get 429 queue_full. Python thread pool with retry; cost is $0.60 at medium.
- Transparent AI image: PNG or WebP, not JPEG, on GPT Image 2.5
For a transparent AI image, request png or webp with background transparent on GPT Image 2.5. JPEG has no alpha. Code to request and verify the alpha channel.
- Trim an AI-generated clip without a media import
Generated artifacts already live on media.sume.com, which is the host video-trim accepts. Use media-imports only for clips from outside Sume.
- Trim and caption a 30-second AI clip: two fixed prices, $0.22
Trim is $0.02 per job and captions are $0.20 per accepted job for videos up to 60 s, so a 30-second AI clip costs $0.22 after the render.
Written by Sume