Number every Shorts series episode on screen with caption cues
YouTube added Shorts series with seasons and episodes in September 2026, per air.io. Burn Episode N onto each clip with Sume caption cues, in Python.

To put a visible episode number on every clip in a Shorts series, send each finished video to Sume's caption endpoint with one authored cue, for example Episode 3, and a start and end time. Because the cue is authored text, no speech-to-text runs, and a clip with no speech still works.
The context: air.io's month-by-month YouTube summary, read on 2026-10-02, says September 2026 brought Shorts series with seasons and episodes. An on-screen number is for viewers who find a single clip outside the series. Sume facts are from the video captions docs.
How do authored cues work?
POST /v1/video-captions takes a public HTTPS video_url and, instead of transcribing, can take cues (or segments): objects with text, start and end in seconds. The docs say authored copy skips speech-to-text and burns exactly that text at those times. script_text, words, cues and segments are mutually exclusive.
Send an Idempotency-Key so a retried request does not queue a second render.
| Field | Value for this job | Why |
|---|---|---|
video_url | Public HTTPS URL of the episode | Required |
cues | One cue: Episode 3, start 0, end 2.5 | Skips speech-to-text |
style | slam or another named style | Omit it and the wording decides |
Idempotency-Key | series-s1-e03 | Safe retries |
What does the Python look like?
The script loops over a list of episodes and submits one caption job each. It prints the HTTP status and body so you can see the job id; poll the job the way the docs describe.
import os
import requests
URL = "https://api.sume.com/v1/video-captions"
HEADERS = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
episodes = [
(1, "https://example.com/s1e01.mp4"),
(2, "https://example.com/s1e02.mp4"),
]
for number, video_url in episodes:
body = {
"video_url": video_url,
"cues": [{"text": f"Episode {number}", "start": 0, "end": 2.5}],
}
headers = {**HEADERS, "Idempotency-Key": f"series-s1-e{number:02d}"}
response = requests.post(URL, json=body, headers=headers, timeout=60)
print(number, response.status_code, response.text[:200])Where should the tag sit?
Placement is a design choice. The design.placement.anchor_ratio field sets the line centre as a fraction of frame height, so you can move the tag away from the area the app's own interface covers. Look at a frame before you render the whole season.
What does this not do?
It does not create the series in YouTube, set the season, or publish. Those steps stay in YouTube Studio or the YouTube API, and the air.io summary is the only source here for what the series feature includes. Replace the example URLs with files you host publicly.
Sources
Related posts
More in Developers
- Shotstack render statuses vs Sume job statuses: a map
Shotstack renders go queued, fetching, rendering, saving, done or failed. Sume jobs go queued, processing, completed, failed or canceled. Map them in code.
- Should you wait for Veo 4? Build on a swappable model id instead
No Veo 4 date, id or price page has been reported. How to ship video work now with the model id in config, and what a swap needs.
- Six scene clips in one agent turn: script_run with a paid-call cap
Fan out six scene generations in one script_run on Sume's MCP, cap them with max_paid_calls, then wait on the child jobs. Sketch, limits and failure behavior.
- Cartesia sonic-3.6-2026-08-27 snapshot: which id Sume accepts
Cartesia's dated snapshot ids never change, but Sume's TTS Router lists only sonic-3.6, 3.5, 3, latest and preview. Here is what that means for repeat takes.
Written by Sume