Word timestamps to video frame numbers at 29.97 fps in Python
Sume STT words[] carry start and end in seconds. Convert them to frame indexes with exact 30000/1001 math so cuts do not drift on long timelines.

Sume STT 1.0 returns words[] as { word, start, end } in seconds from the beginning of the audio. To place a cut or a text overlay on a video timeline, multiply by the frame rate and floor. For NTSC footage use the exact rate 30000/1001, not 29.97, and do the multiplication with fractions.Fraction so rounding does not creep in. The word fields come from the STT result shape in the API reference, read 2026-10-06.
Why 29.97 as a float drifts
29.97 is 2997/100, while the real rate is 30000/1001 (about 29.970030). The gap is 0.00003 frames per frame. Across a one-hour timeline that is about a tenth of a frame at the end, which is small until you round at every step. The bigger error is converting a float second such as 0.12 (which cannot be stored exactly) straight to a frame. Going through Fraction(str(seconds)) takes the decimal the API sent literally.
| Word start (s) | 24 fps | 25 fps | 30000/1001 fps |
|---|---|---|---|
| 0.12 | 2 | 3 | 3 |
| 0.52 | 12 | 13 | 15 |
| 1.00 | 24 | 25 | 29 |
| 60.00 | 1440 | 1500 | 1798 |
The conversion
Floor gives the frame that is on screen when the word begins. Use the same function for the end time if you want the first frame after the word to be the cut point; add one if you need the last frame to include it.
from fractions import Fraction
FPS = Fraction(30000, 1001)
def to_frame(seconds, fps=FPS):
return int(Fraction(str(seconds)) * fps)
def word_frames(words, fps=FPS):
return [
(w["word"], to_frame(w["start"], fps), to_frame(w["end"], fps))
for w in words
]
words = [
{"word": "Hello", "start": 0.12, "end": 0.48},
{"word": "world", "start": 0.52, "end": 1.0},
]
for row in word_frames(words):
print(row)Offsets for chunked audio
If the audio came from several STT calls, add each chunk's start offset to its word times before converting, as in stitching three STT chunks. Convert once, at the end, so the offset and the rate are applied exactly one time each.
Sources
Related posts
More in Developers
- Zed context_servers for the hosted Sume server: no header means OAuth
Add the hosted Sume server to Zed's settings.json context_servers with a url. With no Authorization header Zed runs the MCP OAuth flow, so start with mcp:read.
- Which MCP server lets Claude Code or Cursor generate video and images?
MCP servers that let Claude Code and Cursor make video and images: Sume, fal, Replicate, Runway, Higgsfield. Endpoints, sign-in, billing, setup.
- Idempotency keys for AI video APIs: retry without paying twice
An idempotency key makes a retried create return the original run or job instead of a second paid one. How Sume's Idempotency-Key works on each API.
- Signed webhooks for Sume video runs: events, retries, verification
Sume sends one HMAC-SHA256 signed POST when a Format, Action, or Agent Completion run completes or fails. Verify the raw body and dedupe on request_id.
Written by Sume