YouTube thumbnail test: 3 variants, 2 weeks, desktop only

YouTube's A/B test takes up to three thumbnails and picks by watch time. Generate three distinct 16:9 candidates in one Sume image call and keep them testable.

5 min readSume
All posts

YouTube's A/B test lets you compare up to three thumbnails, titles or both on one video, and it picks the winner by watch time. To feed it, ask Sume's image API for three different 16:9 candidates in one POST /v1/images call with n: 3, then upload the ones you want to test in YouTube Studio.

The test rules below come from YouTube's A/B test titles and thumbnails help page, read on 2026-10-02. Sume's side comes from the Image API docs. Sume does not upload to YouTube or run the test; it only makes the candidates.

How does YouTube's thumbnail test work?

The help page says you can test up to three titles and thumbnails, as title-only, thumbnail-only or combined. The test should finish within two weeks, and it can take a few days depending on impressions and when the video was published. The option with the highest watch time wins; if the result is inconclusive, the first title or combination you uploaded is shown to everyone.

There is a control group that may see only the default version and is left out of the calculation. Results show as Winner, Performed the same or Inconclusive.

A test is also only as good as the traffic behind it. A video with few impressions may not finish in two weeks, and the page itself says it can take a few days or longer depending on impressions and publish time, so run it on videos that already get steady views.

YouTube A/B test rules (read 2026-10-02)
RuleWhat the help page says
VariantsUp to three titles and thumbnails
DurationShould complete within two weeks
WinnerHighest watch time
WhereComputer, in YouTube Studio
Not eligibleShorts, scheduled livestreams, Premieres until they end, Made for Kids, private

Why do the three candidates need to differ?

Watch time can only separate variants that promise different things. Three crops of the same face with a different font colour usually end as Performed the same. Change the idea: one close-up face, one object on a plain colour, one text-led frame. Decide the three concepts first, then write one prompt per concept.

Because the test takes up to two weeks per video, you get few attempts. Spend the effort on the concepts, not on polishing pixels you will throw away.

How do you generate three candidates in one call?

The Image API accepts n for several images per request. The docs say per-model ceilings are lower than the global maximum, so read the n range from the catalog; the catalog example for Seedream 4.5 lists 1 to 4. aspect_ratio: "16:9" is a listed value.

A call blocks for up to 30 seconds and returns 200 with the images, or 202 with a job envelope if it runs long. Check the status code, then poll status_url and result_url for the 202 case. Every n image in a single call shares one prompt, so for three different concepts send three requests with n: 1, or send one request and accept variation.

curl -X POST https://api.sume.com/v1/images \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "bytedance-seed/seedream-4.5",
    "prompt": "YouTube thumbnail: close-up of a surprised cook holding a cracked pan, bold plain background, no text",
    "aspect_ratio": "16:9",
    "n": 3
  }'

What does Sume not do for this test?

It does not upload thumbnails, start the test or read the results. The help page says the feature is on computers in YouTube Studio, so a script cannot drive it from Sume. It also does not guarantee readable text: generated lettering can warp, so a common route is a text-free image and your own title overlay, or a second edit pass as described in thumbnail variants in one call.

Check each image against YouTube's thumbnail requirements before uploading, and keep your own copy: Sume mirrors outputs to media.sume.com, but you should store the files you ship. Shorts are excluded from the test, so for Shorts you compare by publishing, not through this feature.

What is a sensible workflow for a channel?

Pick the video, write three concept prompts, generate one or two images per concept, choose the best of each, upload the three, and set the test. When it ends, record the winning concept, not just the winning file, and write the next round's prompts from it. If you want to refresh an old catalog the same way, see refresh old thumbnails.

Cost is per image and varies by model; read the pricing lines from GET /v1/images/models/{model_id}/endpoints before a batch.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume