Transcribe audio from a private bucket: Sume STT needs a public URL
Sume STT takes a public HTTPS audio_url. If your files sit in a private bucket, here is how to hand them over without opening the whole bucket.

You cannot point Sume STT at a private bucket path. The audio_url field must be a public HTTPS URL, and Sume's docs say signed or private URLs are rejected for media inputs. So the working pattern is to put the file somewhere Sume can fetch it over plain HTTPS, run the job, and take the file down afterwards.
Below is what the schema says, what to do about a private archive, and a short script that submits and polls.
What does the STT schema require?
In the Sume OpenAPI spec the STT 1.0 request has one required field, audio_url, described as a public HTTPS audio URL, with a preference for a Sume media or attachment URL. The asset library workflow states the same rule for other media inputs.
| Location | Works as audio_url? |
|---|---|
| Public HTTPS link | Yes |
| Sume media or attachment URL | Yes, preferred |
| Signed or private URL | Rejected |
| Localhost or a file path | No, it is not a public HTTPS link |
How do you hand over a private file?
Pick the least exposure that works for your data. None of these is something Sume does for you.
- Copy only the files you need to a separate public location, name them with random strings, and delete them when the job is terminal.
- Use a Sume media or attachment URL where your workflow already uploads through Sume.
- Do not publish anything you are not allowed to expose; if the audio is confidential, a public link is the wrong tool.
What does the submit-and-poll code look like?
The key is read from the environment. The job id comes back in data.job.id, and the status route is /v1/jobs/{id}/status.
import os, time, requests
H = {"Authorization": "Bearer " + os.environ["SUME_API_KEY"]}
B = "https://api.sume.com"
r = requests.post(B + "/v1/stt-1.0/transcribe", headers=H, json={
"audio_url": os.environ["AUDIO_URL"],
"language_code": "en",
}, timeout=60)
r.raise_for_status()
job_id = r.json()["data"]["job"]["id"]
while True:
s = requests.get(B + f"/v1/jobs/{job_id}/status", headers=H, timeout=30)
d = s.json()["data"]
if d["terminal"]:
break
time.sleep(d.get("recommended_poll_interval_seconds", 3))
res = requests.get(B + f"/v1/jobs/{job_id}/result", headers=H, timeout=30)
print(d["sume_status"], res.status_code)
What does it cost?
STT is $0.01 per audio minute. Set duration_seconds (up to 600) when you know the length; if you omit it, the reservation is for one minute. For long archives see detaching and transcribing a podcast back catalog.
Sources
Related posts
More in Developers
- Transcribe two minutes of a long video: audio detach range, then STT
Streaming transcribers charge by the hour; you may only need one segment. Detach a range as 16 kHz mono wav, then run one STT job. Caps and codes included.
- 150-language subtitles: which scripts Sume captions document
A translation model can output 150 languages, but Sume documents Latin and Hangul caption styles. Test other scripts on a short clip before a batch.
- Trigger.dev Node 21 warning: which Node runs the Sume SDK
Trigger.dev v4.6.1 added Node.js 21 deprecation warnings. The Sume TypeScript SDK needs Node 18 or later, so tasks on Node 22 or newer are fine.
- Trigger.dev public tokens: keep the Sume key server-side
Trigger.dev v4.6.2 hardened authorization for public tokens. Whatever token your browser holds, a Sume API key must never be one of them. Here is the split.
Written by Sume