Text-to-speech API with curl and jq: one shell script to an MP3
Call the Sume TTS Router from a shell: submit with curl, loop on the status URL with jq until terminal, then download the audio artifact to voiceover.mp3.

You can generate speech from a terminal with three curl calls and jq: submit the job, poll the status URL until terminal is true, then download the audio artifact. The script below does it end to end for the Sume TTS Router and leaves voiceover.mp3 in the current directory. It needs curl, jq, SUME_API_KEY and an avatar handle whose voice is ready.
People ask for this form after every new TTS launch, including Microsoft's MAI-Voice-2.1 on 2026-10-01, because a shell script is the fastest way to judge an API: does it answer in one call, or does it need an SDK and a session? Sume answers with a job. The first response is a small envelope with two URLs, not audio.
Keep the script in version control next to the text it voices, and set the idempotency key from the script name and a version number. Then a rerun after a network drop returns the finished job and does not bill the same characters twice, and a deliberate change gets a new key.
The script
set -euo pipefail
H=(-H "Authorization: Bearer $SUME_API_KEY" -H "Content-Type: application/json")
BODY=$(jq -n --arg h "$SUME_AVATAR_HANDLE" '{model:"sonic-3.6",
transcript:"Spring sale starts Friday. Free shipping on every order.",
avatar_handle:$h, language:"en"}')
JOB=$(curl -sf "${H[@]}" -H "Idempotency-Key: spring-sale-v1" \
https://api.sume.com/v1/tts-router/generate -d "$BODY")
STATUS=$(jq -r .data.status_url <<<"$JOB")
RESULT=$(jq -r .data.result_url <<<"$JOB")
until [ "$(curl -sf "${H[@]}" "$STATUS" | jq -r .data.terminal)" = true ]; do
sleep 3
done
URL=$(curl -sf "${H[@]}" "$RESULT" |
jq -r '.data.result.artifacts[] | select(.type=="audio") | .url')
curl -sfo voiceover.mp3 "$URL"
ls -l voiceover.mp3Why the flags matter
set -e plus curl -f is doing real work here. A failed job never reaches a result: the result call returns 409 job_not_completed, curl -f exits non-zero and the script stops. Replace -f with -sS while debugging and you will see the error body, including the stable error.code.
The Idempotency-Key header is the guard against double spend when you rerun the script. The same key with the same body returns the same job. Change the transcript and keep the key and you get a 409 idempotency_conflict instead of silent reuse, so derive the key from the content when you script a batch.
Fields you will change first
The router needs exactly one text input, transcript or transcript_source, and one voice selector: avatar_id, avatar_handle or voice.id. List avatars with GET /v1/avatar-1.0/avatars and pick one whose voice.status is ready. model must be a catalog id; sonic-latest is an alias that resolves to sonic-3.6.
Audio comes back as MP3 at 44100 Hz and 128 kbps unless you set output_format. A phone system wants { "container": "wav", "sample_rate": 8000, "encoding": "pcm_mulaw" }; an editor wants 48000 Hz WAV.
| Transcript length | Price at $0.0475 per 1,000 characters | Note |
|---|---|---|
| 56 characters (the script above) | $0.0027 | Sume rounds each job up to a whole cent, so $0.01; check the usage field |
| 1,000 characters | $0.0475 | Spaces and punctuation count |
| 20,000 characters | $0.95 | Largest single request |
Scaling the script up
The pricing page quotes $0.0475 per 1,000 characters for TTS (read 2026-10-06). The submit response carries a usage block with the billed amount, so your script can log it with jq .data.usage.
For a long script, split it into sentence-sized or paragraph-sized jobs, give each its own idempotency key and run them in parallel. A single request tops out at 20,000 characters. The envelope fields are documented in Jobs and results.
Two small additions make the script safer to leave in a cron job. Print the job id (jq -r .data.request_id <<<"$JOB") before the loop, so a timeout on your side still leaves you something to look up later; the job keeps running on Sume's side, and you read it from the status URL, you never resubmit. And cap the loop with a counter, for example 120 iterations of 3 seconds, so a stuck network call cannot hold a runner forever. A TTS job that stays queued is rare, but an unbounded until is the kind of line that costs a CI minute bill.
Sources
Related posts
More in Developers
- Sume Timeline output: H.264, AAC 192k and 1-second keyframes
What a Timeline 1.0 MP4 contains for a Reel, Short or TikTok upload: libx264 CRF 20, yuv420p, AAC 192k, faststart, 1-second keyframes, and what you cannot set.
- Timeline 400 on codec, crf or ffmpeg_args: what a Reel spec can send
Sume Timeline refuses codec, crf, preset, filter_complex and ffmpeg_args with a 400. Which output fields it does accept for a 9:16 Reel or Short and why.
- Timeline segment_source_short_looped vs padded: loop or freeze a clip
When a clip is shorter than its slot, Sume Timeline either loops it or holds its last frame. How render.pad_mode auto, loop and freeze decide, with thresholds.
- Timeline mode sync for a 30-second Short: handle both 200 and 202
Sume Timeline mode sync waits up to 30 seconds and returns 200 if the render finished, else 202. Why a Short client must handle both, with a Python poll loop.
Written by Sume