Transcribe a 15-second voice note in one request: STT sync mode
Sume STT can answer in the same HTTP call: mode sync with wait_timeout_seconds up to 30 returns 200, otherwise 202 and a poll. A $0.01 minimum, with curl.

For a 15-second voice note, send POST /v1/stt-1.0/transcribe with mode: "sync" and wait_timeout_seconds up to 30. If the job finishes inside the wait you get a 200 with the finished job; if not, you get a 202 and poll GET /v1/jobs/:id/status and /result. The catalog lists STT at $0.01 per audio minute, with a one-minute reserve when you omit duration_seconds, which is 1 cent.
How sync mode behaves
Sume's communication fields are shared across its generation surfaces. mode is one of async, sync, subscribe or webhook, and wait_timeout_seconds is an integer from 0 to 30. The Timeline docs state the pattern: the default is async; to wait up to 30 seconds for a 200 finished job you pass mode: "sync"; otherwise you get 202 and must poll. Treat a 202 as normal, not as a failure: a sync request is a hint, and a long or queued job will outlive the wait.
| Field | Value | Effect |
|---|---|---|
| audio_url | public HTTPS URL | Required |
| mode | sync | Wait for a finished job up to the timeout |
| wait_timeout_seconds | 25 | 0 to 30; stay under your HTTP client's timeout |
| duration_seconds | 15 | 1 to 600; sets the reserve instead of the 1-minute default |
| language_code | en | 2 to 16 characters; a hint, optional |
A request with a fallback
Give the HTTP client a timeout above the wait, so the client does not give up first, and send an Idempotency-Key so that a retry after a dropped connection returns the same job instead of a second charge.
curl -sS -X POST https://api.sume.com/v1/stt-1.0/transcribe \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: voice-note-0417" \
--max-time 35 \
-d '{
"audio_url": "https://example.com/notes/0417.m4a",
"duration_seconds": 15,
"mode": "sync",
"wait_timeout_seconds": 25
}'
# 200: read the result. 202: poll GET /v1/jobs/{id}/status, then /result.What to read in the result
Four things a short-note handler should check.
textfor the transcript, and the language fields when the job returns them.words[]with{ word, start, end }in seconds from the start of the audio, for captions or trimming.- Nothing about speakers:
diarizeandtag_audio_eventsare fixed on the server, and sending them returns 400. - Money: the job reserves the estimate at submit and captures on completion; failed jobs are refunded.
Handling the 202 path
A 202 means the job is accepted and still running, so store the job id and poll GET /v1/jobs/:id/status every second or two until it reports a finished state, then read /result. Keep the same Idempotency-Key for the retry of the submit, never for a new note: reusing a key on different audio would return the first note's job.
If you would rather not poll, set mode to webhook and verify the signature on the callback. A verifier must refuse an empty secret, so fail closed when the environment variable is missing.
When sync is the wrong shape
Sync suits a note you can answer inside a request. Do not use it for a one-hour file or for a queue of notes: submit those async or with a webhook and let your own worker poll. A 30-second ceiling is a convenience for short clips, not a way to make long jobs fast.
Sources
Related posts
More in Developers
- Transcribe the first 60 seconds of 500 support recordings
500 screen-recorded support sessions: detach seconds 0 to 60 ($5.00), then STT 60 s each ($5.00). Total $10.00 at $0.01 per job and per minute.
- Trim 12 Shorts from one recording with asyncio and a semaphore
Cut 12 Shorts from one recording with asyncio and httpx: 12 trims at $0.02 is $0.24. A semaphore of 4 caps in-flight requests; the limit is your choice.
- TTS 1.0 rejects a model field: use the TTS router to pick sonic-3.6
Sume TTS 1.0 does not accept model. To choose an engine such as sonic-3.6 use /v1/tts-router/generate. Price is the same $0.0475 per 1,000 characters.
- Sume TTS pace test: 1,000 characters a minute decides the cap
Sume TTS stops at 20,000 characters or 1,200 seconds. The break-even is 1,000 characters a minute; measure your voice before a long read. Max $0.95 a job.
Written by Sume