Transcribe audio with curl and jq: a Sume STT shell script
A 16-line bash script that submits audio to Sume STT, polls the job with curl, and prints every word with start and end times through jq. One cent per minute.

To transcribe an audio file from the shell, send one authenticated POST to https://api.sume.com/v1/stt-1.0/transcribe with a public HTTPS audio_url, then poll the job until it is completed and read text and words[] from the result. Sume STT 1.0 is $0.01 per audio minute, so a 2 minute clip is two cents. The script needs only bash, curl and jq, so it fits a cron job, a CI step or a quick check from a terminal.
The endpoint is a job API, not a streaming one: you get a job id back and ask for the result when it is ready. That is the right shape for recorded files, and it keeps the client small.
The request and the result
The body needs audio_url. Optional fields are language_code (a hint such as en or ko; omit it for auto-detect), duration_seconds (1 to 600, which lets Sume reserve the right amount; omit it and Sume reserves one minute), segmentation for sentence rows, and metadata, which is stored with your job and never sent to the provider. Send an Idempotency-Key header so a retry does not create a second paid job.
A submit returns 202 with request_id, which is the job id. Poll GET /v1/jobs/{id}/status with a pause between reads, and stop on completed, failed or canceled. Then GET /v1/jobs/{id}/result returns text, language_code, words[] with word, start and end in seconds from the audio start, and segments[] if you asked for them.
| Step | Call | In the script |
|---|---|---|
| Submit | POST /v1/stt-1.0/transcribe | api -X POST with jq-built body |
| Wait | GET /v1/jobs/{id}/status | while loop with case |
| Read | GET /v1/jobs/{id}/result | jq prints start, end, word |
The script
Run SUME_API_KEY=... ./stt.sh https://media.sume.com/artifacts/artf_demo/clip.wav stt-001. The output is tab-separated, so it pipes straight into awk or a spreadsheet.
#!/usr/bin/env bash
set -euo pipefail
AUDIO_URL="$1"; KEY="$2"; BASE=https://api.sume.com
api() { curl -fsS -H "Authorization: Bearer $SUME_API_KEY" -H "Content-Type: application/json" "$@"; }
ID=$(api -X POST "$BASE/v1/stt-1.0/transcribe" -H "Idempotency-Key: $KEY" \
-d "$(jq -n --arg u "$AUDIO_URL" '{audio_url: $u, duration_seconds: 120}')" | jq -r .request_id)
while :; do
STATUS=$(api "$BASE/v1/jobs/$ID/status" | jq -r .status)
case "$STATUS" in
completed) break ;;
failed | canceled) echo "job $ID ended $STATUS" >&2; exit 1 ;;
esac
sleep 2
done
api "$BASE/v1/jobs/$ID/result" |
jq -r '.words[] | select(.type == "word" or .type == null) | "\(.start)\t\(.end)\t\(.word)"'the shell details worth knowing
set -euo pipefail makes the script stop at the first failing command, and curl -fsS turns an HTTP error into a non-zero exit with a short message. Without -f, curl exits 0 on a 402 or 409 and jq then reads an error body as if it were a job.
jq -n --arg u "$AUDIO_URL" builds the JSON body, so a URL with special characters is escaped correctly. Avoid assembling JSON by string concatenation in shell.
The case statement stops the loop on failed and canceled and prints the job id to standard error, so a calling script can read it. The final jq filter keeps entries whose type is word or missing, which drops the spacing tokens that STT results can include.
To make this a webhook instead of a poll, add "mode": "webhook" and a public HTTPS webhook_url to the body. Sume sends a single terminal event, job.completed, job.failed or job.canceled.
Limits and prices
One job takes at most 10 minutes of audio. For longer recordings, cut the audio into slices of up to 600 seconds and send one job per slice, adding each slice's start offset to its word times when you merge. The audio must be at a public HTTPS URL, and Sume media URLs are the preferred source. If the audio is inside a video, audio detach extracts a 16 kHz mono WAV for $0.01 per job.
If the request is refused for balance, the API returns 402; a changed body under a reused idempotency key returns 409; a rate limit returns 429. Do not resubmit a paid request just because your own process timed out, because the job may still be running. Read the status first, as the jobs docs advise.
Keep the returned job id next to your own record of the file. If a result looks wrong later, that id is what lets you fetch the same job again without paying for a second transcription, and the metadata object you sent at submit time is stored with it, so you can tag each job with your own file or ticket id and find it again.
Sources
Related posts
More in Developers
- unsupported_capability names sume/auto: fix a Sume ad clip request
A 400 unsupported_capability on sume/auto hides the resolved model but lists accepted values in supported. How to fix duration, resolution or audio.
- verifyWebhook in a fetch handler: four rules, 204 for unknown events
Use @sume-com/sdk verifyWebhook on the raw body, await it, treat false as 401 and answer unknown events with 204. A runnable handler for Workers, Deno and Node.
- Wan 3.0 API request cheat sheet: three modes, 2 to 30 seconds
Wan 3.0 on Sume: the request body for text, first/last frame and reference modes, the 480p/720p/1080p rates and the 2 to 30 second window, on one page.
- GPT Image 1 to GPT Image 2.5 on Sume: what changes in the output
Moving from GPT Image 1 to ChatGPT Image 2.5 on Sume changes the response (URL, not base64), default quality, size grid and failures.
Written by Sume