AI thumbnail variants in one call: n=4, pick one, then polish
Ask for four thumbnail options in one POST /v1/images call with n up to 4, pick the best by eye, then run a polish edit. Cost per step and when n is a waste.
Use n: 4 with aspect_ratio: "16:9" to get four thumbnail options from one request, choose one by eye, then edit that single image to polish text or color. The Image API accepts n up to 10, but each model sets its own ceiling in the catalog (the docs example, Seedream 4.5, lists 1 to 4), so read the n range for your model from GET /v1/images/models before relying on four (Image API docs).
This beats four separate calls because you send one request, one response lists four URLs, and usage.cost shows the total in a single field.
Plan the two steps
Each step has a different job.
| Step | Settings | Goal |
|---|---|---|
| Options | n: 4, aspect_ratio: 16:9, quality: low or medium | Find a composition |
| Polish | n: 1, picked URL in input_references, quality: high | Fix text, color, crop |
Code for both steps
The script prints four URLs, asks for a choice and runs the polish edit.
import os
import requests
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
U = "https://api.sume.com/v1/images"
M = "openai/gpt-image-2.5"
a = requests.post(U, headers=H, timeout=90, json={"model": M, "n": 4,
"aspect_ratio": "16:9", "quality": "medium",
"prompt": "YouTube thumbnail, surprised cook holding a giant pancake"})
a.raise_for_status()
urls = [d["url"] for d in a.json()["data"]]
for i, u in enumerate(urls):
print(i, u)
b = requests.post(U, headers=H, timeout=90, json={"model": M,
"aspect_ratio": "auto", "quality": "high",
"prompt": "Same thumbnail, brighter colors, keep the composition",
"input_references": [{"type": "image_url", "image_url": {
"url": urls[int(input("pick: "))]}}]})
print(b.status_code, b.json())When is n=4 wasteful?
If you already have a layout, skip the options and just edit. If text accuracy is the goal, a wide choice of compositions does not help, so move to a higher quality or a model with better text rendering. Four images cost four times one, and fal lists GPT Image 2.5 at $0.0060 for a 1024 square at low (read 2026-10-01, fal); Sume's own price is in the response.
What to ask the polish step for
Polish prompts work best when they change one thing: brighten the colors, make the headline larger, or move the subject. Name the thing to keep ('keep the face and the pancake') so the edit does not become a new image. If the polish drifts too far, go back to the picked option and try a second, smaller edit.
Limits
A 4-image call may exceed the 30-second sync wait and return a 202 job, in which case you poll the job, not the original call. Thumbnails shrink a lot, so view the pick at small size before polishing. YouTube's own size rules were not checked for this post.
Sources
Related posts
More in Use cases
- AI-written script on a public-interest topic: Article 50 text label
Article 50(4) text labels cover published public-interest text with no human review or editorial control. What counts as review, per the Commission.
- Amazon DVA+ rolls out from late October: get your video assets ready
Amazon says DVA+ begins rolling out in late October and advertisers approve creatives before launch. A checklist for preparing holiday video assets with Sume.
- Storyboard animatic from stills with a Sume timeline
Check pacing and shot order before paying for video: render storyboard stills as a silent Timeline 1.0 animatic, with an unbilled plan call first.
- AI Act marking exemption for B2B and industrial output: how narrow
The Commission FAQ says a narrow Article 50(2) marking exemption is envisaged for B2B or industrial outputs, with conditions in the guidelines. What it lists.
Written by Sume