Reddit 15-second engaged views: does your voiceover CTA land?

Reddit's Engaged Video Views beta bills videos over 15 s at 15 s. Use TTS word timestamps to check that your call to action is spoken before second 15.

5 min readSume
All posts

If you buy Reddit video ads on the new 15-second engaged-view metric, the first 15 seconds are the ones that count. The October platform update from Orthotropy (read 2026-10-06) says the Engaged Video Views public beta started on 2026-09-01 and that videos longer than 15 seconds bill at 15. So a voiceover whose call to action arrives at second 17 is spending the budget on the lead-in. You can check that before you render anything: ask Sume TTS for word timestamps, find the call-to-action word and compare its start time to 15 seconds. The script does exactly that.

This is a pre-flight check, not a claim about how Reddit counts a view. The source says what it says; read Reddit's own ad documentation before you plan spend, and treat the 15-second line below as your own budget for where the offer must be spoken.

Find the CTA word and its start time

import json, os, time, urllib.request as u
KEY = os.environ["SUME_API_KEY"]

def call(url, body=None, key=None):
    h = {"Authorization": "Bearer " + KEY, "Content-Type": "application/json"}
    if key: h["Idempotency-Key"] = key
    data = json.dumps(body).encode() if body else None
    return json.load(u.urlopen(u.Request(url, data=data, headers=h)))["data"]

LIMIT, CTA = 15.0, "code"
SCRIPT = ("Tired of slow mornings? Meet the grinder that is ready in ten seconds. "
          "Order this week and use code BREW for free shipping.")
job = call("https://api.sume.com/v1/tts-router/generate", {
    "model": "sonic-3.6", "language": "en", "avatar_handle": os.environ["SUME_AVATAR_HANDLE"],
    "transcript": SCRIPT, "timestamps": {"words": True}}, "reddit-15s-v1")
while not job["terminal"]:
    time.sleep(job.get("next_poll_after_seconds") or 2)
    job = call(job["status_url"])
words = call(job["result_url"])["result"]["words"]
hit = next((w for w in words if w["word"].strip(" ,.!?").lower() == CTA), None)
total = max(w["end"] for w in words)
if hit is None:
    print("CTA word not found")
elif hit["start"] < LIMIT:
    print(f"CTA at {hit['start']:.1f}s of {total:.1f}s: inside the first {LIMIT:.0f} s")
else:
    print(f"CTA at {hit['start']:.1f}s: move the offer up or raise speed (0.6 to 1.5)")

Reading the result

The words come back as word, start and end in seconds, ordered by start time. The script strips punctuation and compares lowercase, so "code" matches "code,". If your CTA is a phrase, search for the first word of it. Change LIMIT if the platform changes the rule. If the offer lands late, there are two fixes: cut the lead-in, or raise speed, which is allowed from 0.6 to 1.5, and measure again. Cutting is better; a faster read to save two seconds is a change in tone that listeners notice.

Cost of the check

A TTS check costs a fraction of a cent: the script above is 123 characters, or about $0.006 at $0.0475 per 1,000 (pricing, read 2026-10-06). Run it on every variant of the script before you pay for video, because a failed pre-flight is a cheaper way to learn that the offer arrives late. The rest of the job envelope, including terminal and next_poll_after_seconds, is in Jobs and results.

Where the offer lands against a 15 s window (read 2026-10-06)
Spoken offer starts atInside the first 15 sWhat to do
9.8 sYesShip the script
14.6 sBarelyTrim the lead-in so a slower read still fits
17.2 sNoCut the first sentence or restructure

Other controls that move the timing

Speed is not the only dial. The TTS request schema accepts speed from 0.6 to 1.5, volume from 0.5 to 2 and a free-text emotion of up to 64 characters in its generation config. Speed changes the total length and so shifts every word time; emotion can change pacing too, which is why the check should be re-run whenever you change either one rather than once per script.

Ask for segmentation with mode: "sentence" alongside the word timestamps if you also want the sentence boundaries. The API then returns gapless segments, with each segment starting where the previous one ends and a 70 ms post-word boundary by default, which makes it easy to see which sentence the call to action belongs to and how long the lead-in sentences run before it.

Test hook variants

Run the check on every variant you plan to test. If you try three hooks, that is three short TTS jobs, and the one whose offer lands earliest is the one to put into video first. Keep the spoken script and the on-screen text in step: if the caption says the code at second 8 and the voice says it at second 12, viewers see a mismatch. Word timestamps give you the same numbers to place a caption or text card exactly.

Reuse the timings

The same check works for other platforms with a hard early window. Pair it with the length fit in fit a voiceover to a 30-second slot when the whole ad also has a maximum, and with the platform sizes in Reddit video ad specs. Word timings are the common thread: once you have them for a script, they drive captions, trims and checks like this one.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume