YouTube Shorts 2x playback: will your burned-in captions still read?
Viewers can double the speed of a Short. Burned-in captions play at that speed too. Halve each phrase's time, then tune Sume caption phrasing to fit.

Yes, they play at 2x with the rest of the video, so a caption that sits on screen for one second now sits there for half a second. YouTube's Shorts update post says viewers "can now double the playback speed on a Short" (read 2026-10-04). Burned-in captions are pixels in the video, so nothing re-times them for a fast viewer. Check how long each phrase stays up, halve it, and loosen the phrasing for any card that gets too short.
This is arithmetic, not a platform rule. YouTube publishes no minimum reading time for captions, and neither does Sume. The threshold in the script below is yours to set.
What changed on the viewer side
The same June 25, 2026 post lists a Clear screen mode, a mute control by tapping the screen, an adjustable Shorts timer, a heart in place of the thumbs-up, and the dislike button being phased out. The speed control is the one that touches caption timing, because it changes how long every on-screen word lasts.
Text you burned into the picture has no timing of its own. The timing you authored is the timing at 1x, and half of it at 2x.
| Phrase on screen at 1x | At 2x | Our editorial read for a three-word card |
|---|---|---|
| 0.8 s | 0.4 s | Hard to catch at either speed |
| 1.2 s | 0.6 s | Fine at 1x, tight at 2x |
| 2.0 s | 1.0 s | Comfortable at both |
Measure your cards before you ship
If you author phrase cards yourself, you already have the times. Sume's caption job accepts cues with text, start and end in seconds, which skip speech-to-text and burn exactly that copy (video captions docs). The script reads a cues file in that shape, prints the 1x and 2x on-screen time of each card and flags the ones under your minimum.
import json, sys
cues_path = sys.argv[1] # JSON list of {text, start, end}
minimum_at_2x = float(sys.argv[2]) if len(sys.argv) > 2 else 0.6
with open(cues_path) as f:
cues = json.load(f)
short = 0
for cue in cues:
shown = cue["end"] - cue["start"]
at_2x = shown / 2
flag = "" if at_2x >= minimum_at_2x else " <- short at 2x"
short += bool(flag)
print(f"{shown:5.2f}s {at_2x:5.2f}s {cue['text']}{flag}")
print(f"{short} of {len(cues)} cards under {minimum_at_2x}s at 2x")
Tune the phrasing instead of re-timing the video
Cards get short when each one carries few words and the speaker moves quickly. Sume's design.phrasing override has three knobs: max_words, max_chars and pause_seconds. Raising max_words and max_chars should group more words per card, so each card lasts longer; re-run the script on the result to confirm. pause_seconds sets how long a silence has to be before a card breaks. The placement and typography groups let you move the line and size it, for example to keep it clear of the icons YouTube draws.
Two limits apply. design is not supported on the punch and tiktok-green styles, which render on a path that reads none of these tokens. And a number outside its documented range is a 400 at request time, so a bad value fails before it bills. A standalone caption job is $0.20 for videos up to 60 seconds.
- Re-burn with
source_caption_idand a newdesignto try other phrasing without a second transcription. - Keep the script flag at the minimum you would accept on a phone held at arm's length.
- Preview the result at 2x in the YouTube app before you publish.
A worked pass
Say a 30 second Short has 40 cards and the script flags 12 of them under your 0.6 s minimum at 2x. Those 12 are usually the fast stretches, where the speaker packs many short words together. Re-burn once with a higher max_words and max_chars, run the script again, and look at the count. If it fell to 3, you have fixed most of the problem for $0.20. If it did not move, the cause is probably the speech itself, and the honest fix is a script with fewer words in that stretch.
Do this once per style, not once per video. A style's phrasing tokens apply to every clip you caption with the same design, so a good setting from one Short carries to the rest of the batch. Spot-check the first clip of each new speaker, since pace varies by person.
What this does not cover
Sume cannot make a viewer watch at 1x, and it cannot read or change the speed a viewer picked. If your Short depends on fast speech, burned-in cards are only half the answer, since the voice also doubles. For where the captions sit relative to the new heart and mute controls, see the heart and caption placement note.
Sources
Related posts
More in Media tools
- How to assemble a long-form video with the Timeline 1.0 API
Timeline 1.0 renders one audio spine plus 1 to 200 ordered video slots into one MP4. Every URL must be Sume-hosted; the plan preflight is unbilled.
- How to burn captions onto a video with the Sume API
Send a public HTTPS video URL to POST /v1/video-captions and get a job-backed captioned video, timed by speech-to-text or by text you supply.
- How to extract frames from a video with the Sume API
POST /v1/video-frames returns stills at the times you name from one Sume-hosted clip, as durable images at source size. The call is unbilled.
- How to use Sume's Timeline compose and Timeline audio APIs
Timeline compose puts one still and one video in the same frame as a new MP4. Timeline audio joins or splits Sume-hosted audio into reusable files.
Written by Sume