Thumbnail A/B test: five GPT Image 2.5 variants at medium vs high

Five 16:9 thumbnail variants cost about $0.04 at medium and $0.16 at high in GPT Image 2.5 output tokens. A budget for 20 videos and the review step.

5 min readSume
All posts

Five 16:9 thumbnail variants at 1536x864 cost about $0.042 in GPT Image 2.5 output tokens at medium and $0.162 at high, before input tokens and Sume pricing. For twenty videos that is about $0.84 at medium for 100 candidates. The cheap way to run an A/B test is to generate all variants at medium, pick two, and rerun only those at high.

The budget

Per-image estimates come from the token formula behind Fal's GPT Image 2.5 Flare page and the Image API docs: $30 per million output tokens. 1536x864 is 16:9, both edges are multiples of 16, and it passes the size rules in OpenAI's image generation guide.

GPT Image 2.5 output estimate at 1536x864 (read 2026-10-04)
QualityPer image5 variants100 variants (20 videos)
low$0.0036$0.018$0.36
medium$0.0084$0.042$0.84
high$0.03234$0.1617$3.23
xhigh$0.05751$0.28755$5.75

A workflow that spends little

Do not generate five variants at high. Generate them at medium, put them in front of whoever decides, and promote only the best two. A reference photo of the presenter's face is input tokens on top; at $8 per million per Fal's page it is a small line, but it is not zero, and the count is an estimate.

  • Round 1: five variants at medium, one prompt per hook idea. The catalog lists n up to 4 for openai/gpt-image-2.5, so send n: 4 plus a second request of n: 1, or five single-image requests.
  • Review: cut to two by a fixed rubric (face visible, three-word headline, contrast).
  • Round 2: rerun the two prompts at high, n: 2 each.
  • Ship: pick, then add any exact text in your editor if the model got a letter wrong.

Text on thumbnails

OpenAI's image generation guide says text rendering and composition precision remain areas for improvement, so plan to check every headline letter by eye. A variant with a misspelled word is a failed variant, however good the image looks. Keep the text to two or three words and consider adding it in a design tool after generation.

curl -X POST "https://api.sume.com/v1/images" \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "openai/gpt-image-2.5",
    "prompt": "YouTube thumbnail, shocked presenter on the left, bold headline NEW PHONE on the right",
    "aspect_ratio": "16:9",
    "quality": "medium",
    "n": 4
  }'

What to check

Read the n range on the model's catalog record before sending a count: the docs say per-model ceilings are lower than the route's 1 to 10, and the GPT Image 2.5 record lists 4, which is why the sample sends 4 and a fifth image goes in a second request. Large n is one of the configurations that tends to degrade to a 202 job, so handle that as described in Jobs and results. Also read usage.cost once to confirm the real billed amount before you scale to 100.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume