Sketch plus face photo to a thumbnail with GPT Image 2.5
Send a rough layout sketch and a portrait as two references and ask GPT Image 2.5 for a 16:9 thumbnail. The prompt, a 1280x720 Sume request and what to verify.
To turn a thumbnail sketch and a portrait into a finished image with GPT Image 2.5, send both as input_references, tell the prompt that Image 1 is the layout sketch and Image 2 is the face to keep, and request a 16:9 size such as 1280x720. Sume has no separate Sketch endpoint, and the sketch is simply another reference image.
This post relies on OpenAI's Image prompting guide, fal's GPT Image 2.5 guide, both read on 2026-10-02, and Sume's Image API page. Only use a portrait of someone who agreed to appear.
What does the guidance say about sketches?
OpenAI's guide lists sketches to realistic renders as a technique: keep the layout, and add photorealistic materials and lighting in the prompt. Its role-assignment advice, identify each input by number and purpose, is what lets one request carry both a sketch and a face.
How do I split the roles?
The sketch decides where things sit. The portrait decides who it is. The prompt decides everything else. Say each in one sentence so they do not compete.
| Input | Role | Prompt sentence |
|---|---|---|
| Image 1 | Layout sketch | Use Image 1 only for composition: keep the positions and sizes of the boxes. |
| Image 2 | Portrait | Image 2 is the person: keep the face and identity exactly. |
| Text | Headline | The big text reads "I quit my job" in double quotes, bold, right side. |
| Style | Finish | Photorealistic, high contrast, warm rim light. |
What does the Sume request look like?
1280x720 is a valid custom size for GPT Image 2.5 on Sume: both edges are multiples of 16, the ratio is 16:9 and the pixel count is above the 655,360 minimum. With two references, set the size explicitly instead of relying on auto. Set n above 1 only if you want to compare options; each image is billed.
curl -X POST "https://api.sume.com/v1/images" \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "openai/gpt-image-2.5",
"prompt": "Image 1 is a layout sketch: use it only for composition, keep the positions and sizes of its boxes. Image 2 is the person: keep her face and identity exactly. Render a photorealistic YouTube thumbnail, bold headline \"I quit my job\" on the right, warm rim light.",
"image_size": "1280x720",
"quality": "high",
"input_references": [
{"type": "image_url", "image_url": {"url": "https://example.com/sketch.png"}},
{"type": "image_url", "image_url": {"url": "https://example.com/portrait.jpg"}}
]
}'What should I check in the result?
Compare the output with the sketch for box positions and with the portrait for the face. Read the headline letter by letter. If the layout drifted, say so in the prompt (for example, name the exact side the text sits on) and shorten the style sentence. If the face drifted, restate the identity constraint, as covered in the same-face prompt block.
Sume does not guarantee that a sketch is followed pixel for pixel, and the platform you upload to sets its own thumbnail size and file limits; check those before exporting. For choosing among several options in one call, see AI thumbnail variants in one call.
How should I draw the sketch?
Keep it simple: boxes for the face and the text, a rough arrow or object shape, and nothing else. Label each box in the prompt, for example "left box is the face, right box is the headline". Dark lines on a white page are easier for the model to read than a photograph of a notebook.
Make the sketch the same shape as the output, 16:9 here, so the boxes do not need to be reinterpreted. If the first result misses the layout, fix the sketch before you rewrite the prompt.
Sources
- Image API
- Image prompting (OpenAI, read 2026-10-02)
- [How To Use GPT Image 2.5: Prompts & Workflows [2026] (fal, read 2026-10-02)](https://fal.ai/learn/tools/how-to-use-gpt-image-2-5)
Related posts
More in Use cases
- Not signing the EU transparency code: what you must show instead
The Commission says non-signatories must show their Article 50 measures are adequate, may face more information requests, and are assessed by each authority.
- Steam Next Fest October 2026: cut a demo trailer with captions
Steam Next Fest runs 19 to 26 October 2026. Cut a 30-second demo trailer from gameplay captures using trims, Timeline and caption cues, for under a dollar.
- Synthesia Sessions alternative: avatar briefing clips by API
Synthesia launched Sessions for live avatar roleplay. Sume renders scripted avatar clips instead: what that covers in sales training, and what it does not.
- Thanksgiving recipe video: step cards with caption cues
Make a 40-second Thanksgiving recipe video from five phone shots: join them in Timeline 1.0, burn numbered step cards with caption cues, add a music bed.
Written by Sume