Clickable transcript from Sume STT word timestamps in Python
Turn a Sume speech-to-text result into HTML where each word seeks the audio player to its start time. Runnable Python, with the result envelope handled safely.

Sume's speech-to-text always returns word timings, so a clickable transcript needs no extra request. Each word has word, start and end, in seconds. Wrap each word in a span that carries its start time, then let a few lines of script seek an audio element when a word is clicked.
Submit and fetch
Send POST /v1/stt-1.0/transcribe with a public HTTPS audio_url, and duration_seconds between 1 and 600 so the reservation matches your clip. The price is $0.01 per audio minute. Poll GET /v1/jobs/:id/status, and call GET /v1/jobs/:id/result once completed; before that it returns 409 job_not_completed. See jobs and results.
The renderer
The function below finds the words list wherever it sits in the result JSON, so it does not depend on the envelope layout. Save the result body to result.json first.
import html, json
def find_words(node):
if isinstance(node, dict):
if isinstance(node.get("words"), list):
return node["words"]
for v in node.values():
found = find_words(v)
if found:
return found
return []
def render(words):
spans = [
'<span data-t="%.2f">%s</span>' % (w.get("start", 0), html.escape(w["word"]))
for w in words if "word" in w
]
js = "document.addEventListener('click',e=>{const t=e.target.dataset.t;" \
"if(t){const a=document.querySelector('audio');a.currentTime=+t;a.play();}})"
return "<audio controls src='clip.mp3'></audio><p>" + " ".join(spans) + "</p><script>" + js + "</script>"
data = json.load(open("result.json"))
open("transcript.html", "w").write(render(find_words(data)))Making it better
- Highlight the current word on the audio element's
timeupdateevent by comparingcurrentTimewith eachstart. - Group words into sentences with
segmentation: {mode: sentence}on the request, and render one paragraph per segment. - Host the audio on a stable URL;
audio_urlmust be public HTTPS, and keep it reachable until the job finishes.
Limits
Word times are those the transcription returned, so a mis-heard word is clickable but wrong. Speaker labels are not an option in this API. Transcripts are capped at 10 minutes per request, so a longer file needs splitting; see Timeline audio for a split. Cue-style output is covered in subtitle cues from sentence segments.
Related posts
More in Developers
- Sume job error category quota or queue vs 402 and 429
A Sume job error category quota means add funds or lower cost; queue means retry with the same key. They sit on the job, apart from 402 and 429 at submit.
- Sume job failed with worker_timeout: poll again or retry?
A Sume job error in the worker_timeout or generation_timeout category means poll status or retry later. runtime_unavailable means retry later, gently.
- Sume job metadata on Kling motion control and H3 Max lip sync
Both Kling 3.0 Motion Control and MiniMax H3 Max Lip Sync accept a metadata object stored with the Sume job request. It is not sent to the provider.
- jobs_result with job_ids: read ok per entry, re-read failed_job_ids
A Sume MCP jobs_result batch returns ok plus value or error per id. One job_not_completed does not fail the rest; re-read partial_failure.failed_job_ids.
Written by Sume