Sume STT metadata: tag a transcript job with your own ids
The metadata field on a Sume STT request is stored with the job and is not sent to the speech provider. Use it to tie transcripts to your records.

Put your own ids in the metadata object of the Sume STT request. The OpenAPI spec describes it as optional caller metadata stored with the Sume job request, and not sent to the provider. That makes it the right place for an episode id, a ticket number or a batch label, and the wrong place for anything you expect the transcription model to see.
It is useful because job ids come from Sume while the records you care about come from you. Metadata is how the two meet.
What does the spec guarantee?
In the Sume OpenAPI spec the STT request lists metadata as optional, stored with the job request, and not forwarded to the provider. The spec does not promise how it is echoed on every read, so confirm on a test job where it appears before you build a reconciliation on it.
| Field | Purpose | Sent to the provider? |
|---|---|---|
| audio_url | The recording to transcribe | Needed to fetch the audio |
| language_code | Language hint, omit for auto-detect | Used for transcription |
| metadata | Your own tags | No |
| webhook_url | Terminal callback target | No, it is Sume delivery |
What should you put in it?
Keep it small and boring. Use identifiers, not content.
- Your own record id, such as an episode or call id.
- A batch label so you can find a run later.
- A schema version for your own parser.
- Not secrets and not personal data you would not want stored.
How do you send it?
A submit call with tags:
import os, requests
H = {"Authorization": "Bearer " + os.environ["SUME_API_KEY"]}
r = requests.post("https://api.sume.com/v1/stt-1.0/transcribe",
headers=H, timeout=60,
json={
"audio_url": os.environ["AUDIO_URL"],
"duration_seconds": 240,
"metadata": {"episode_id": "ep-0142", "batch": "oct-archive"},
})
r.raise_for_status()
print(r.json()["data"]["job"]["id"])
Is there a related tag for end users?
Yes, the same idea applies across Sume jobs; see tagging voice and image jobs with an end-user id. For batches of recordings, pair it with the pattern in a podcast back catalog.
Sources
Related posts
More in Developers
- Sume STT sentence segmentation fails closed when no words are timed
If the speech provider returns no timed words, a Sume STT request with segmentation returns a typed error instead of guessed sentences. Plan for it.
- Test a Sume STT webhook locally: webhook_url must be public HTTPS
Sume rejects localhost, private-network and non-HTTPS webhook_url values. Put a tunnel in front of your dev server, or poll while you build.
- Authenticate the Sume CLI on a CI runner without a browser login
On CI, skip sume login: install the CLI, run sume auth setup with an API key from a secret, and confirm with sume auth status before any job step.
- sume/auto for a former Sora feature: when to pin a model
Sume's sume/auto picks a family and never says which. Good for general clips, wrong when a brand needs one look. How to choose between auto and a pinned id.
Written by Sume