Real time speech to text API: what a file-based API can do

Sume's speech to text API isn't real time: it transcribes recordings of up to 10 minutes at a URL. Chunked recordings give near-live transcripts.

5 min readSume
All posts

A real time speech to text API takes live audio, such as a microphone or a phone call, while it is being spoken and returns text within the conversation. Sume's speech to text API is not real time: STT 1.0 transcribes a recording at a public HTTPS URL, up to 10 minutes per request, and returns the transcript when the job finishes. It has no microphone or live-stream input and no WebSocket.

The facts below come from the STT 1.0 schema in the Sume API reference (the OpenAPI document behind the API reference docs) and Jobs and results, read on 2026-09-29. If you need words on screen as someone talks, you need a streaming speech service; if a delay of one chunk is acceptable, the pattern below works with a file-based API.

What does Sume's speech to text API take and return?

One request, one recording. The body has no field for audio bytes or a stream, only a URL.

From the STT 1.0 request and result schemas in the Sume API reference, read 2026-09-29.
ItemWhat the API reference says
audio_urlRequired; a public HTTPS audio URL
duration_secondsOptional, 1–600, used to reserve usage; omit it and 1 minute is reserved
language_codeOptional hint; omit it for auto-detect
Resulttext, language fields when available, and words[]; in current code each word carries start and end in seconds when they are returned
SpeakersNo speaker labels; diarize is fixed server-side and can't be sent

Is there a WebSocket or streaming endpoint?

No. The docs say there is no SSE or WebSocket transport on the Developer API today, and GET /v1/jobs/:id/events is a pull snapshot, not a stream. mode: "sync" holds the submit request for at most 30 seconds; it isn't a stream either. Job webhooks are terminal-only: job.completed, job.failed or job.canceled, with no progress or partial callbacks. So there are no partial transcripts: you get the whole chunk's text at once.

How do I get near-live transcripts from a file-based API?

Cut the live audio into short recordings and transcribe each one as soon as it closes. The transcript trails the speaker by at least one chunk plus the job's run time, which the docs don't state, so it fits notes and after-the-fact captions, not live conversation.

  • Record fixed-length chunks, for example 60 seconds each, on your side.
  • Put each closed chunk at a public HTTPS URL in your own storage. Sume has no public route for uploading local files.
  • Submit POST /v1/stt-1.0/transcribe with the chunk's audio_url, its duration_seconds, and a webhook_url, so the transcript arrives without a poll loop. Keep polling status_url as a backup.
  • Add each chunk's start time to its word times: they are seconds from the start of that chunk's audio.
  • Expect split words at chunk edges: a word that straddles a boundary is cut between two recordings. Overlap chunks slightly and drop duplicate words if that matters.
curl -X POST https://api.sume.com/v1/stt-1.0/transcribe \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: call-7-chunk-0012" \
  -d '{
    "audio_url": "https://example.com/live/call-7/chunk-0012.m4a",
    "duration_seconds": 60,
    "language_code": "en",
    "webhook_url": "https://example.com/webhooks/sume"
  }'

What does it cost?

STT 1.0 costs $0.01 per audio minute, plus a 5.5% agent fee by default. At that rate, 60 minutes of audio comes to $0.60 before the fee. Send duration_seconds with each chunk: without it, the request reserves 1 minute. Speech to text in Python walks through one full request and result.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume