Clickable transcript from Sume STT word timestamps in Python

Turn a Sume speech-to-text result into HTML where each word seeks the audio player to its start time. Runnable Python, with the result envelope handled safely.

5 min readSume
All posts

Sume's speech-to-text always returns word timings, so a clickable transcript needs no extra request. Each word has word, start and end, in seconds. Wrap each word in a span that carries its start time, then let a few lines of script seek an audio element when a word is clicked.

Submit and fetch

Send POST /v1/stt-1.0/transcribe with a public HTTPS audio_url, and duration_seconds between 1 and 600 so the reservation matches your clip. The price is $0.01 per audio minute. Poll GET /v1/jobs/:id/status, and call GET /v1/jobs/:id/result once completed; before that it returns 409 job_not_completed. See jobs and results.

The renderer

The function below finds the words list wherever it sits in the result JSON, so it does not depend on the envelope layout. Save the result body to result.json first.

import html, json

def find_words(node):
    if isinstance(node, dict):
        if isinstance(node.get("words"), list):
            return node["words"]
        for v in node.values():
            found = find_words(v)
            if found:
                return found
    return []

def render(words):
    spans = [
        '<span data-t="%.2f">%s</span>' % (w.get("start", 0), html.escape(w["word"]))
        for w in words if "word" in w
    ]
    js = "document.addEventListener('click',e=>{const t=e.target.dataset.t;" \
         "if(t){const a=document.querySelector('audio');a.currentTime=+t;a.play();}})"
    return "<audio controls src='clip.mp3'></audio><p>" + " ".join(spans) + "</p><script>" + js + "</script>"

data = json.load(open("result.json"))
open("transcript.html", "w").write(render(find_words(data)))

Making it better

  • Highlight the current word on the audio element's timeupdate event by comparing currentTime with each start.
  • Group words into sentences with segmentation: {mode: sentence} on the request, and render one paragraph per segment.
  • Host the audio on a stable URL; audio_url must be public HTTPS, and keep it reachable until the job finishes.

Limits

Word times are those the transcription returned, so a mis-heard word is clickable but wrong. Speaker labels are not an option in this API. Transcripts are capped at 10 minutes per request, so a longer file needs splitting; see Timeline audio for a split. Cue-style output is covered in subtitle cues from sentence segments.

Related posts

More in Developers

All Developers posts

Written by Sume