Speech-to-text for short voice notes: one request, sync mode

Send a voice note's public URL to Sume STT with mode sync and a 30-second wait. A cent per minute, plus the poll fallback when the job outlasts the wait.

5 min readSume
All posts

For a voice note under a minute or two, submit POST /v1/stt-1.0/transcribe with mode: "sync" and wait_timeout_seconds: 30. Sume holds the HTTP request open for up to 30 seconds, and a short clip usually finishes inside that window. A 45-second note costs one cent. If the job is still running when the wait ends, the response is still a success and you poll the status URL instead of resubmitting.

The code below always reads the result through the job endpoints, so it behaves the same whether the wait finished the job or not.

What sync mode does and does not do

sync and subscribe are aliases for a bounded wait, clamped to 0 to 30 seconds. The wait bounds the HTTP request, not the job. If the job is queued or the waiter budget is full, the response carries a sync.timed_out or sync.capacity_exhausted marker and you keep polling status_url; do not submit a second paid job for the same note.

Authentication is exactly one of Authorization: Bearer <key> or x-api-key: <key>; sending both is rejected.

Sume STT 1.0 for voice notes, per-request facts (read 2026-10-08)
ItemValue
RoutePOST /v1/stt-1.0/transcribe
Inputaudio_url, public HTTPS
Max audio per job600 seconds
Reservation hintduration_seconds, 1 to 600; omitted means 1 minute
Wait in sync modewait_timeout_seconds 0 to 30
PriceAbout $0.01 per audio minute, rounded up; 1 cent minimum
Result fieldstext, language_code, language_probability, words

The call

Voice notes arrive in many containers. The Sume docs show .m4a and .wav examples, so test your own format on one note before you automate. The audio must be at a public HTTPS URL; a signed URL that expires in minutes works if the job fetches it in time.

import os, time, requests

B = "https://api.sume.com"
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}

def voice_note(url, seconds=60):
    r = requests.post(f"{B}/v1/stt-1.0/transcribe", headers=H,
        json={"audio_url": url, "duration_seconds": seconds,
              "mode": "sync", "wait_timeout_seconds": 30})
    r.raise_for_status()
    job = r.json()["data"]["request_id"]
    while True:
        s = requests.get(f"{B}/v1/jobs/{job}/status", headers=H).json()["data"]
        if s["status"] in ("FAILED", "CANCELED"):
            raise RuntimeError(s["status"])
        if s["result_ready"]:
            break
        time.sleep(s.get("next_poll_after_seconds") or 3)
    res = requests.get(f"{B}/v1/jobs/{job}/result", headers=H)
    return res.json()["data"]["result"]["text"]

Limits

There is no streaming transcription: Sume returns text after the file is processed, not while someone talks. STT 1.0 has no speaker labels. Use language_code as a hint when you know the language; omit it to auto-detect, and read language_probability before trusting a short note in a mixed-language workplace.

Operational notes

Send an Idempotency-Key per voice note if your client can retry, so a network blip does not create a second paid job. Store the Sume job id with the note id; the status and result endpoints both key on it.

Handle error responses by retrying with the same key. For a burst of notes, submit them asynchronously and collect results through a webhook instead of holding a request open for each one. Sync mode is a convenience for a single interactive request, not a way to run a queue.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume