Whistle word timestamps to burned captions with Sume's words field
Turn Whistle's on-device word timings into burned-in captions: map them to words with text, start and end, then call Sume video captions with no STT run.

To caption a video with Whistle's output, convert each word to {text, start, end} in seconds and send the list as words to POST /v1/video-captions together with a public HTTPS video_url. When you send words, Sume does not run speech-to-text; it burns your text at your times. The conversion is a few lines of code, and the one trap is the clock offset when you transcribe in 30-second windows.
Cactus Compute says Whistle returns word-level timestamps with start, end and probability scores, and takes at most 30 seconds of 16 kHz mono audio per pass (Whistle launch post, read 2026-10-11). The field names in your Whistle binding may differ from the ones below, so adapt the mapping to what your build returns.
What shape does the caption job want?
The Sume caption docs describe words as word-level items with text, start and end in seconds, and allow only one of script_text, words, cues and segments per request. style is optional: punch is one of the named styles, and Latin text works on the Latin display styles. Korean text must use a Hangul style instead.
The video must be a public HTTPS URL that Sume can fetch. A standalone caption job reserves and captures $0.20 for a video up to 60 seconds under the current estimate; read GET /v1/catalog for the live figure.
How do you map the words?
The function below drops empty tokens and zero-length words, rounds the times, and adds an offset so that the second 30-second window starts at 30.0 instead of 0.0. The submit function refuses an empty API key before it sends anything, then posts with an Idempotency-Key so a retry cannot create a second paid job. Run it as is to see the mapping printed.
import os, requests
def to_caption_words(whistle_words, offset=0.0):
out = []
for w in whistle_words: # items with word, start, end
text = str(w["word"]).strip()
if text and w["end"] > w["start"]:
out.append({"text": text,
"start": round(w["start"] + offset, 3),
"end": round(w["end"] + offset, 3)})
return out
def burn(video_url, words, key):
if not key:
raise RuntimeError("SUME_API_KEY is empty")
r = requests.post(
"https://api.sume.com/v1/video-captions",
headers={"Authorization": f"Bearer {key}",
"Idempotency-Key": "whistle-captions-001"},
json={"video_url": video_url, "style": "punch", "words": words},
timeout=60)
r.raise_for_status()
return r.json()
if __name__ == "__main__":
demo = [{"word": "Hello", "start": 0.1, "end": 0.4},
{"word": "Sume", "start": 0.5, "end": 0.9}]
print(to_caption_words(demo))
What can go wrong?
Three things come up in practice.
- Offsets: if you transcribed in windows, add window_index times 30 seconds to every word in that window, or the captions restart at zero.
- Overlap: if you overlapped windows, remove duplicated words at the seam before you send the list.
- Wording: Whistle supports keyword biasing for phrases, but if the spelling of a product name still comes out wrong, send script_text instead, which keeps STT timings and aligns your approved text to them.
Where does the result appear?
Restyling is cheap in effort: send source_caption_id with a new style instead of video_url, and Sume reuses the source video and the word timings it already holds. That keeps your Whistle timings in play without uploading them again.
The job is asynchronous. Keep the id from the response and poll GET /v1/jobs/{id}/status with backoff, then read /result or the video-captions resource for the captioned video URL. Do not resubmit the paid request because a local timeout expired.
Sources
Related posts
More in Media tools
- Open and close a clip with a white flash: the fade filter
Fade a clip in from white and out to white with two fade filters in Sume video filter. Out time comes from the clip duration; audio is untouched.
- Add a zoomed detail inset to a video: crop, scale, overlay
Show a magnified product detail in the corner of the same clip using one Sume video-filter graph: split, crop, scale 2.4x, white border, overlay. $0.02.
- How to assemble a long-form video with the Timeline 1.0 API
Timeline 1.0 renders one audio spine plus 1 to 200 ordered video slots into one MP4. Every URL must be Sume-hosted; the plan preflight is unbilled.
- How to burn captions onto a video with the Sume API
Send a public HTTPS video URL to POST /v1/video-captions and get a job-backed captioned video, timed by speech-to-text or by text you supply.
Written by Sume