Whistle word timestamps to burned captions with Sume's words field

Turn Whistle's on-device word timings into burned-in captions: map them to words with text, start and end, then call Sume video captions with no STT run.

5 min readSume
All posts

To caption a video with Whistle's output, convert each word to {text, start, end} in seconds and send the list as words to POST /v1/video-captions together with a public HTTPS video_url. When you send words, Sume does not run speech-to-text; it burns your text at your times. The conversion is a few lines of code, and the one trap is the clock offset when you transcribe in 30-second windows.

Cactus Compute says Whistle returns word-level timestamps with start, end and probability scores, and takes at most 30 seconds of 16 kHz mono audio per pass (Whistle launch post, read 2026-10-11). The field names in your Whistle binding may differ from the ones below, so adapt the mapping to what your build returns.

What shape does the caption job want?

The Sume caption docs describe words as word-level items with text, start and end in seconds, and allow only one of script_text, words, cues and segments per request. style is optional: punch is one of the named styles, and Latin text works on the Latin display styles. Korean text must use a Hangul style instead.

The video must be a public HTTPS URL that Sume can fetch. A standalone caption job reserves and captures $0.20 for a video up to 60 seconds under the current estimate; read GET /v1/catalog for the live figure.

How do you map the words?

The function below drops empty tokens and zero-length words, rounds the times, and adds an offset so that the second 30-second window starts at 30.0 instead of 0.0. The submit function refuses an empty API key before it sends anything, then posts with an Idempotency-Key so a retry cannot create a second paid job. Run it as is to see the mapping printed.

import os, requests

def to_caption_words(whistle_words, offset=0.0):
    out = []
    for w in whistle_words:  # items with word, start, end
        text = str(w["word"]).strip()
        if text and w["end"] > w["start"]:
            out.append({"text": text,
                        "start": round(w["start"] + offset, 3),
                        "end": round(w["end"] + offset, 3)})
    return out

def burn(video_url, words, key):
    if not key:
        raise RuntimeError("SUME_API_KEY is empty")
    r = requests.post(
        "https://api.sume.com/v1/video-captions",
        headers={"Authorization": f"Bearer {key}",
                 "Idempotency-Key": "whistle-captions-001"},
        json={"video_url": video_url, "style": "punch", "words": words},
        timeout=60)
    r.raise_for_status()
    return r.json()

if __name__ == "__main__":
    demo = [{"word": "Hello", "start": 0.1, "end": 0.4},
            {"word": "Sume", "start": 0.5, "end": 0.9}]
    print(to_caption_words(demo))

What can go wrong?

Three things come up in practice.

  • Offsets: if you transcribed in windows, add window_index times 30 seconds to every word in that window, or the captions restart at zero.
  • Overlap: if you overlapped windows, remove duplicated words at the seam before you send the list.
  • Wording: Whistle supports keyword biasing for phrases, but if the spelling of a product name still comes out wrong, send script_text instead, which keeps STT timings and aligns your approved text to them.

Where does the result appear?

Restyling is cheap in effort: send source_caption_id with a new style instead of video_url, and Sume reuses the source video and the word timings it already holds. That keeps your Whistle timings in play without uploading them again.

The job is asynchronous. Keep the id from the response and poll GET /v1/jobs/{id}/status with backoff, then read /result or the video-captions resource for the captioned video URL. Do not resubmit the paid request because a local timeout expired.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume