YouTube dynamic thumbnails: generate three options by API
Dynamic thumbnails need three candidate files. Send three Sume image jobs in webhook mode, each with its own Idempotency-Key, and collect the files when done.
YouTube Studio's dynamic thumbnails have YouTube recommend the best of three options to different audience segments, so you need three finished thumbnail files. On Sume, send three image jobs, each with its own Idempotency-Key, in webhook mode so every submit returns 202 straight away and you collect the files when each job finishes.
Read 2026-09-30. YouTube's post, as read, says nothing about how candidates are chosen, so this page covers only the file production. For choosing what differs between the three, see YouTube dynamic thumbnails: three for different audiences.
Why three separate jobs?
Each thumbnail has its own prompt, so each is its own request. Reference images go in image_urls (1 to 10 public HTTPS URLs) on Image 1.0. Sending the same references with three prompts keeps the subject constant while the framing or words change.
What mode should I use?
The Jobs and results page says webhook mode returns 202 with the job envelope and polling URLs, and the callback is stored. The job id is in the first response in every mode. Keep polling as a backup, since a delivery can fail.
| Mode | What the submit returns |
|---|---|
async (default) | 202; poll status_url, then read result_url |
sync | Waits up to 30 s; if not terminal, poll, do not resubmit |
webhook | 202; wait for the callback, and keep polling as backup |
What goes in each request?
A prompt, the reference URLs, and aspect_ratio. YouTube thumbnails are 16:9, so set "16:9". On edits, the Image API says to prefer "auto" to match the reference shape; omitting the field is not the same as auto.
curl -X POST https://api.sume.com/v1/image-1.0/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: thumb-option-a" \
-d '{
"prompt": "Same subject, bold close crop, empty space on the left",
"image_urls": ["https://media.sume.com/artifacts/artf_demo/frame.png"],
"aspect_ratio": "16:9",
"mode": "webhook",
"webhook_url": "https://example.com/hooks/sume"
}'What happens after the three finish?
Each completed job returns Sume-hosted images under result.artifacts[]. Download the three, check faces and any lettering against your original, then use them in YouTube Studio. Sume does not upload to YouTube or run the test.
Sources
Related posts
More in Use cases
- YouTube inauthentic content policy and bulk AI video runs
YouTube's inauthentic content policy targets mass-produced, generic AI templates. How a 100-item Format bulk queue can give every video its own input.
- YouTube playlistImages.insert 50MB: cover art from a clip frame
YouTube's API history raised playlistImages.insert from 2MB to 50MB on 2026-09-14. A Sume video-frames still is a way to produce the image to send.
- YouTube Shorts up to 3 minutes: trim a long video to fit
YouTube's help page says Shorts can run up to three minutes and must be square or vertical. Sume video-trim cuts a range from a longer clip into a new MP4.
- YouTube Shorts series: queue episodes as one Format bulk run
Queue 1 to 100 episodes of a Shorts series in one Sume Format bulk run, set a concurrency of 1 to 16, and branch on counts.failed, not on completed.
Written by Sume