Hormozi-style captions: what they are and how to make them

Hormozi-style captions are big uppercase words shown one at a time, timed to speech, with key words in a contrast color. How to make them by API.

4 min readSume
All posts

Hormozi-style captions are big, uppercase words that pop onto the screen one or two at a time, in sync with the speech, with key words picked out in a contrasting color. The style is named after Alex Hormozi's short-form videos; it is a look anyone can use, and Sume has no connection to him.

To make them automatically you need word-level timings and a caption style that shows one word at a time with keyword accents. Sume's slam caption style is built that way. The details below come from Sume's Video captions docs and, where marked, its caption renderer code, read on 2026-09-28.

What makes a caption Hormozi-style?

  • Very few words on screen: one, sometimes two, replaced as each is spoken.
  • A tall, heavy display font in uppercase, large enough to read on a phone with the sound off.
  • Centered, a little below the middle of a vertical frame, clear of the face and the platform buttons.
  • One accent color on the words that carry the point, such as a number, a verb, or a result.
  • Word-by-word timing, so each word appears as it is said.

How does Sume's slam style match that look?

In current code, slam renders each word on its own, in capitals, in the Anton display face, with gold on the words it picks as keywords:

From Video captions and Sume's caption renderer code, read 2026-09-28.
Elementslam (current code)
Words per card1: one word at a time is the style's identity
Case and faceUppercase, Anton
ColorsWhite words, gold (#FFD700) on keyword picks
PositionCentered, the line at 0.58 of the frame height
Keyword picksWords of 5 characters or fewer that are not stopwords, skipping the first word; if none qualifies, the last word that is not a stopword
FrameA black overlay at 18% opacity over the whole video

How do I add Hormozi-style captions to a video?

Send the video's public HTTPS URL to POST /v1/video-captions and name slam. Name it rather than omitting style: an omitted style does resolve to slam for Latin wording, but the docs say a defaulted render drops the gold accent and keeps the fill color on the spoken word. Naming slam keeps its gold. Without a script, Sume times the words with speech-to-text; send words (text, start, end in seconds) to burn your own copy at your own times.

curl -X POST https://api.sume.com/v1/video-captions \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: slam-captions-001" \
  -d '{
    "video_url": "https://example.com/talking-head.mp4",
    "style": "slam",
    "design": { "colors": { "active": "#22D3EE" } }
  }'

Can I change the colors, position, or words per card?

Colors and position, yes. design.colors.active sets the accent color (cyan in the sample above), and design.placement.anchor_ratio moves the line up or down as a fraction of the frame height. Words per card, no: slam shows one word at a time, and in current code its phrasing caps have no effect. You also can't choose which words get the accent; the pick is a fixed rule in current code. Customize captions lists every token and range.

What are the limits?

  • A Latin display face: Korean copy on slam returns 400 (caption_hangul_text_latin_style).
  • In current code a caption job refuses a video over 60 seconds or one with no audio stream. For longer videos, split, caption, and rejoin.
  • Every word is capitalized in current code, so the burned text will not match your script's casing.
  • Each job reserves a fixed amount for videos up to 60 seconds; confirm the price in GET /v1/catalog.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume