YouTube dynamic thumbnails: make three for different audiences

YouTube lists dynamic thumbnails among new Studio tools. Make three thumbnail candidates for three audiences from one video frame with the Sume image API.

4 min readSume
All posts

To make three thumbnails for different audiences, decide what each audience cares about, then send one image request per audience from the same source frame, changing only the subject emphasis and the words. On Sume that is one POST /v1/images call to openai/gpt-image-2.5 per audience, each with the frame as an input_references entry.

YouTube's Made on YouTube post names "dynamic thumbnails" among its new Studio tools but, as read on 2026-09-29, does not describe how they choose between candidates, so this page does not either. It covers the preparation you control: three candidates, ready to upload or test.

What is YouTube offering?

The post lists dynamic thumbnails next to channel-matched thumbnail generation and an Ask Studio that runs in the background to suggest refreshed thumbnails. It also says creators have run more than 40 million title and thumbnail experiments since 2024. Treat anything beyond those names as not yet published.

What should change between the three thumbnails?

Keep the frame, the face and the size fixed; vary one thing an audience would notice. Three audiences for one tutorial might be beginners, experts and shoppers.

Example audience split for one video; the request fields are from the Image API docs, read 2026-09-29.
AudiencePrompt emphasisRequest
BeginnersRelieved face, one plain word of space for a title1 call
ExpertsResult on screen, no face close-up1 call
ShoppersProduct held toward camera1 call

How do I get the frame and send the requests?

Pull a still from your video with video frames, which takes a media.sume.com clip and returns durable image artifacts, and is unbilled. Reference URLs for the image call must be public HTTPS, so open the frame URL in a logged-out window before you use it. openai/gpt-image-2.5 takes up to 16 references, and image_size of 3840 by 2160 fits its custom-pixel rules: both edges multiples of 16, a maximum edge of 3840, and at most 8,294,400 pixels.

curl -X POST https://api.sume.com/v1/images \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "openai/gpt-image-2.5",
    "prompt": "Same person, relieved smile, bright background, empty space on the left, no text",
    "input_references": [
      { "type": "image_url", "image_url": { "url": "https://media.sume.com/artifacts/artf_demo/frame.jpg" } }
    ],
    "image_size": { "width": 3840, "height": 2160 }
  }'

Which frame should I start from?

Ask the frames call for a handful of moments with at[], which takes 1 to 24 values in seconds, and pick the one with the clearest face and the most open background. Each value must be inside the clip or the worker fails with frame_time_out_of_range. The default jpeg is fine for reference use; png is lossless if you want to inspect detail.

Use the same frame for all three audiences. If the frame changes as well as the prompt, you cannot tell which change made a candidate better.

How do I check the three before publishing?

Each result is a new image, not your frame, so compare faces and any product labels with the original. Set a title in an editor rather than trusting generated lettering. Sume ends at the image file; you upload and test in YouTube Studio.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume