Thumbnail A/B test: five GPT Image 2.5 variants at medium vs high
Five 16:9 thumbnail variants cost about $0.04 at medium and $0.16 at high in GPT Image 2.5 output tokens. A budget for 20 videos and the review step.
Five 16:9 thumbnail variants at 1536x864 cost about $0.042 in GPT Image 2.5 output tokens at medium and $0.162 at high, before input tokens and Sume pricing. For twenty videos that is about $0.84 at medium for 100 candidates. The cheap way to run an A/B test is to generate all variants at medium, pick two, and rerun only those at high.
The budget
Per-image estimates come from the token formula behind Fal's GPT Image 2.5 Flare page and the Image API docs: $30 per million output tokens. 1536x864 is 16:9, both edges are multiples of 16, and it passes the size rules in OpenAI's image generation guide.
| Quality | Per image | 5 variants | 100 variants (20 videos) |
|---|---|---|---|
| low | $0.0036 | $0.018 | $0.36 |
| medium | $0.0084 | $0.042 | $0.84 |
| high | $0.03234 | $0.1617 | $3.23 |
| xhigh | $0.05751 | $0.28755 | $5.75 |
A workflow that spends little
Do not generate five variants at high. Generate them at medium, put them in front of whoever decides, and promote only the best two. A reference photo of the presenter's face is input tokens on top; at $8 per million per Fal's page it is a small line, but it is not zero, and the count is an estimate.
- Round 1: five variants at
medium, one prompt per hook idea. The catalog listsnup to 4 foropenai/gpt-image-2.5, so sendn: 4plus a second request ofn: 1, or five single-image requests. - Review: cut to two by a fixed rubric (face visible, three-word headline, contrast).
- Round 2: rerun the two prompts at
high,n: 2each. - Ship: pick, then add any exact text in your editor if the model got a letter wrong.
Text on thumbnails
OpenAI's image generation guide says text rendering and composition precision remain areas for improvement, so plan to check every headline letter by eye. A variant with a misspelled word is a failed variant, however good the image looks. Keep the text to two or three words and consider adding it in a design tool after generation.
curl -X POST "https://api.sume.com/v1/images" \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "openai/gpt-image-2.5",
"prompt": "YouTube thumbnail, shocked presenter on the left, bold headline NEW PHONE on the right",
"aspect_ratio": "16:9",
"quality": "medium",
"n": 4
}'What to check
Read the n range on the model's catalog record before sending a count: the docs say per-model ceilings are lower than the route's 1 to 10, and the GPT Image 2.5 record lists 4, which is why the sample sends 4 and a fifth image goes in a second request. Large n is one of the configurations that tends to degrade to a 202 job, so handle that as described in Jobs and results. Also read usage.cost once to confirm the real billed amount before you scale to 100.
Sources
Related posts
More in Use cases
- TikTok creator-labeled AI tag: who sets it on an API upload
TikTok's Direct Post API has an is_aigc field for the creator-labeled AI tag. Sume's trim, captions and timeline jobs do not set it; your upload step must.
- TikTok's Q3 2026 ad products: which need new video, which do not
Of nine items in TikTok's Q3 2026 preview, TopView self-serve and TopReach touch creative. A table sorts them and maps the clip work to Sume.
- TikTok Shop Good-tier listings: five images over 600x600, batched
TikTok Shop's listing quality tiers ask for five or more images above 600x600 and a clean main image. How to batch that set on Sume. Read 2026-10-03.
- TikTok Shop LIVE: stills may cover only half the screen
TikTok Shop seller guidance caps stills and slideshows at 50% of the screen during a LIVE. Plan overlay layouts with motion clips from Sume instead of stills.
Written by Sume