Real time speech to text API: what a file-based API can do
Sume's speech to text API isn't real time: it transcribes recordings of up to 10 minutes at a URL. Chunked recordings give near-live transcripts.

A real time speech to text API takes live audio, such as a microphone or a phone call, while it is being spoken and returns text within the conversation. Sume's speech to text API is not real time: STT 1.0 transcribes a recording at a public HTTPS URL, up to 10 minutes per request, and returns the transcript when the job finishes. It has no microphone or live-stream input and no WebSocket.
The facts below come from the STT 1.0 schema in the Sume API reference (the OpenAPI document behind the API reference docs) and Jobs and results, read on 2026-09-29. If you need words on screen as someone talks, you need a streaming speech service; if a delay of one chunk is acceptable, the pattern below works with a file-based API.
What does Sume's speech to text API take and return?
One request, one recording. The body has no field for audio bytes or a stream, only a URL.
| Item | What the API reference says |
|---|---|
audio_url | Required; a public HTTPS audio URL |
duration_seconds | Optional, 1–600, used to reserve usage; omit it and 1 minute is reserved |
language_code | Optional hint; omit it for auto-detect |
| Result | text, language fields when available, and words[]; in current code each word carries start and end in seconds when they are returned |
| Speakers | No speaker labels; diarize is fixed server-side and can't be sent |
Is there a WebSocket or streaming endpoint?
No. The docs say there is no SSE or WebSocket transport on the Developer API today, and GET /v1/jobs/:id/events is a pull snapshot, not a stream. mode: "sync" holds the submit request for at most 30 seconds; it isn't a stream either. Job webhooks are terminal-only: job.completed, job.failed or job.canceled, with no progress or partial callbacks. So there are no partial transcripts: you get the whole chunk's text at once.
How do I get near-live transcripts from a file-based API?
Cut the live audio into short recordings and transcribe each one as soon as it closes. The transcript trails the speaker by at least one chunk plus the job's run time, which the docs don't state, so it fits notes and after-the-fact captions, not live conversation.
- Record fixed-length chunks, for example 60 seconds each, on your side.
- Put each closed chunk at a public HTTPS URL in your own storage. Sume has no public route for uploading local files.
- Submit
POST /v1/stt-1.0/transcribewith the chunk'saudio_url, itsduration_seconds, and awebhook_url, so the transcript arrives without a poll loop. Keep pollingstatus_urlas a backup. - Add each chunk's start time to its word times: they are seconds from the start of that chunk's audio.
- Expect split words at chunk edges: a word that straddles a boundary is cut between two recordings. Overlap chunks slightly and drop duplicate words if that matters.
curl -X POST https://api.sume.com/v1/stt-1.0/transcribe \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: call-7-chunk-0012" \
-d '{
"audio_url": "https://example.com/live/call-7/chunk-0012.m4a",
"duration_seconds": 60,
"language_code": "en",
"webhook_url": "https://example.com/webhooks/sume"
}'What does it cost?
STT 1.0 costs $0.01 per audio minute, plus a 5.5% agent fee by default. At that rate, 60 minutes of audio comes to $0.60 before the fee. Send duration_seconds with each chunk: without it, the request reserves 1 minute. Speech to text in Python walks through one full request and result.
Sources
Related posts
More in Developers
- Remove background from image in Node.js (JavaScript API)
Remove an image background from Node.js: call a background-removal API with fetch on your server, poll the job, and save the transparent PNG.
- Remove background from image in Python with an API
Remove an image background in Python: POST the image URL with requests, poll the job, then save the transparent PNG. A full script and the price.
- Replicate API rate limits: 600 creates a minute, then 429
Replicate's API allows 600 prediction creates and 3,000 other requests per minute. Low credit and no card tighten it; over the limit you get a 429.
- Runway API rate limit: usage tiers, concurrency and 429s
Runway's API has no requests-per-minute limit. Usage tiers cap concurrency per model, generations per 24 hours and monthly spend.
Written by Sume