Spend more, save more: three offer tiers in one clip, one caption job
Show a tiered Black Friday offer as three timed cue cards on one clip. One $0.20 caption job covers all tiers, and the savings are computed, not typed.

A spend-more-save-more offer fits in one short clip as three timed cue cards, one per tier, burned in a single caption job. The job is $0.20 for a clip up to 60 seconds, however many cards you send, so a three-tier offer costs the same as a one-line one. The clean master is untouched, so a fourth tier or a changed threshold later is one more job on the same master.
One job, many cards
The captions API takes cues (also called segments): objects with text, start and end in seconds. Sume burns that text at those times and does not run speech-to-text, which is why it works on a silent product clip. You can send only one of script_text, words, cues and segments in a request.
Compute the savings
A tier card has two numbers a customer will check: the threshold and what they save. Typing both invites a mismatch ("Spend $100, save 15% ($20 off)"). Derive the dollar line from the percent, as in the table, and burn the generated string.
Two cautions. A percentage saved at the threshold is the saving on that exact basket, so the card should say "at $150" if the dollar figure assumes it. And if the percentage applies only to some items, say so in the card rather than implying a store-wide discount.
| Spend | Percent off | Saving at that spend | Cue window (s) |
|---|---|---|---|
| $50 | 10% | $5.00 | 0.5-3.0 |
| $100 | 15% | $15.00 | 3.0-5.5 |
| $150 | 20% | $30.00 | 5.5-8.0 |
The request
The script prints the request body for POST /v1/video-captions. Add an Idempotency-Key header when you send it. The result is a captioned MP4. If tiers change, the clean master is unchanged and you re-run with new cues.
import json
tiers = [(50, 10), (100, 15), (150, 20)] # spend, percent off
cues = []
for i, (spend, pct) in enumerate(tiers):
save = spend * pct / 100
cues.append({"text": f"Spend ${spend}, save {pct}% (${save:.2f} off)",
"start": 0.5 + i * 2.5, "end": 3.0 + i * 2.5})
body = {"video_url": "https://example.com/hero.mp4", "style": "slam", "cues": cues}
print(json.dumps(body, indent=1))Pacing
Give each card about 2.5 seconds, which is the time to read a short line twice on a phone. If you need more than three tiers, make a second clip rather than shrinking the time per card. Pull frames in the middle of each window to check that no card is cut by the next one. Also check the last tier lasts long enough to read before the clip ends, and that the product is still visible behind the text, since a card that hides the thing it sells is a poor trade.
When not to use this
If the offer has conditions that need more than one line (excluded brands, a cap on the discount), a video card is the wrong place to carry them in full. Point to the page with the terms and keep the clip to the headline. Regulators in several markets treat omitted conditions as misleading, so check your own.
Sources
Related posts
More in Use cases
- Steam's AI disclosure: do trailer and capsule art made with AI count?
Steam's content survey asks about pre-generated and live-generated AI. Which one covers a trailer or capsule art made with a Sume model, and what must match.
- Stepping away from a TikTok Shop LIVE: Pause LIVE, not a loop
TikTok Shop's page says to use Pause LIVE when you step away and bars AI or pre-recorded audio. Where a Sume clip belongs, and where it does not, around a LIVE.
- Store-wide sale claim in a video: choose the wording from data
ACCC sweeps flagged misleading site-wide and store-wide claims. Pick the card wording from the catalog data, then burn it with caption cues. Read 2026-10-05.
- Subscription box reveal from six item photos: Wan references, $1.25
Wan 3.0 takes up to 10 reference images: six item photos make a 10 s box reveal at 720p for $1.25 on Sume. References guide the look; they are not exact frames.
Written by Sume