Word timestamps to video frame numbers at 29.97 fps in Python

Sume STT words[] carry start and end in seconds. Convert them to frame indexes with exact 30000/1001 math so cuts do not drift on long timelines.

5 min readSume
All posts

Sume STT 1.0 returns words[] as { word, start, end } in seconds from the beginning of the audio. To place a cut or a text overlay on a video timeline, multiply by the frame rate and floor. For NTSC footage use the exact rate 30000/1001, not 29.97, and do the multiplication with fractions.Fraction so rounding does not creep in. The word fields come from the STT result shape in the API reference, read 2026-10-06.

Why 29.97 as a float drifts

29.97 is 2997/100, while the real rate is 30000/1001 (about 29.970030). The gap is 0.00003 frames per frame. Across a one-hour timeline that is about a tenth of a frame at the end, which is small until you round at every step. The bigger error is converting a float second such as 0.12 (which cannot be stored exactly) straight to a frame. Going through Fraction(str(seconds)) takes the decimal the API sent literally.

Frame index for the same second values at two frame rates, computed with the code below.
Word start (s)24 fps25 fps30000/1001 fps
0.12233
0.52121315
1.00242529
60.00144015001798

The conversion

Floor gives the frame that is on screen when the word begins. Use the same function for the end time if you want the first frame after the word to be the cut point; add one if you need the last frame to include it.

from fractions import Fraction

FPS = Fraction(30000, 1001)


def to_frame(seconds, fps=FPS):
    return int(Fraction(str(seconds)) * fps)


def word_frames(words, fps=FPS):
    return [
        (w["word"], to_frame(w["start"], fps), to_frame(w["end"], fps))
        for w in words
    ]


words = [
    {"word": "Hello", "start": 0.12, "end": 0.48},
    {"word": "world", "start": 0.52, "end": 1.0},
]
for row in word_frames(words):
    print(row)

Offsets for chunked audio

If the audio came from several STT calls, add each chunk's start offset to its word times before converting, as in stitching three STT chunks. Convert once, at the end, so the offset and the rate are applied exactly one time each.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume