How-to Reel: burn numbered step captions with authored cues

A 45-second tutorial Reel with Step 1 to Step 5 on screen: pass authored cues to Sume video captions, skip speech-to-text, and pay $0.20 for up to 60 seconds.

5 min readSume
All posts

To put Step 1 to Step 5 on a tutorial Reel, send Sume's video captions endpoint a list of authored cues, each with text, start and end in seconds. Authored cues skip speech-to-text entirely, so they work on a silent screen recording, and the job is a fixed $0.20 for videos up to 60 seconds under the current estimate.

Instagram's Reels guide (read 2026-10-03) names 15 to 60 seconds as the sweet spot for tutorials. It also says Instagram transcribes spoken audio automatically and lets you review and style the captions. Step labels are not speech, so they are the part you add yourself.

What the endpoint accepts

POST /v1/video-captions requires video_url, a fetchable public HTTPS video. For overlay copy pass cues (or segments) with text, start and end. cues, segments, words and script_text are mutually exclusive. A silent clip given no cues fails caption_no_speech with next_action: use_overlay_captions, which is the signal to do exactly this. SRT uploads are not supported; phrase-level text goes in as cues.

Cue plan for a 45-second how-to Reel (read 2026-10-03)
CueTextstart (s)end (s)
1Step 1: Open Settings1.08.5
2Step 2: Tap Notifications9.017.0
3Step 3: Turn on Reels alerts17.527.0
4Step 4: Choose who can message you27.537.0
5Step 5: Save37.544.0

The call

The request needs an Idempotency-Key so a retry does not bill twice. The script below posts the cues and prints the job id; poll GET /v1/jobs/{id}/status and read /result as described in Jobs and results.

import os, requests

cues = [
    {"text": "Step 1: Open Settings", "start": 1.0, "end": 8.5},
    {"text": "Step 2: Tap Notifications", "start": 9.0, "end": 17.0},
    {"text": "Step 3: Turn on Reels alerts", "start": 17.5, "end": 27.0},
    {"text": "Step 4: Choose who can message you", "start": 27.5, "end": 37.0},
    {"text": "Step 5: Save", "start": 37.5, "end": 44.0},
]
r = requests.post(
    "https://api.sume.com/v1/video-captions",
    headers={
        "Authorization": "Bearer " + os.environ["SUME_API_KEY"],
        "Idempotency-Key": "howto-steps-001",
    },
    json={"video_url": os.environ["CLIP_URL"], "style": "slam", "cues": cues},
)
r.raise_for_status()
print(r.json())

Look and legibility

The style field takes slam, punch or tiktok-green for Latin text; slam is the default for Latin when you omit it. A design object can tune colours, typography, placement and phrasing on styles that support it, but the docs say design is not supported on punch or tiktok-green. Keep cue times from overlapping, which is the simplest rule to reason about.

What Sume does not do

Sume does not watch the screen recording and place Step labels where the action happens; you choose the times, ideally by sampling stills with video inspect. It does not make the clip itself, and the $0.20 figure is the fixed estimate for videos up to 60 seconds, so confirm a longer clip in GET /v1/catalog before you submit. If your Reel is spoken, drop cues and pass nothing or a script_text to correct names instead.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume