Sketch plus face photo to a thumbnail with GPT Image 2.5

Send a rough layout sketch and a portrait as two references and ask GPT Image 2.5 for a 16:9 thumbnail. The prompt, a 1280x720 Sume request and what to verify.

5 min readSume
All posts

To turn a thumbnail sketch and a portrait into a finished image with GPT Image 2.5, send both as input_references, tell the prompt that Image 1 is the layout sketch and Image 2 is the face to keep, and request a 16:9 size such as 1280x720. Sume has no separate Sketch endpoint, and the sketch is simply another reference image.

This post relies on OpenAI's Image prompting guide, fal's GPT Image 2.5 guide, both read on 2026-10-02, and Sume's Image API page. Only use a portrait of someone who agreed to appear.

What does the guidance say about sketches?

OpenAI's guide lists sketches to realistic renders as a technique: keep the layout, and add photorealistic materials and lighting in the prompt. Its role-assignment advice, identify each input by number and purpose, is what lets one request carry both a sketch and a face.

How do I split the roles?

The sketch decides where things sit. The portrait decides who it is. The prompt decides everything else. Say each in one sentence so they do not compete.

Roles in a sketch plus portrait request (OpenAI guidance, read 2026-10-02)
InputRolePrompt sentence
Image 1Layout sketchUse Image 1 only for composition: keep the positions and sizes of the boxes.
Image 2PortraitImage 2 is the person: keep the face and identity exactly.
TextHeadlineThe big text reads "I quit my job" in double quotes, bold, right side.
StyleFinishPhotorealistic, high contrast, warm rim light.

What does the Sume request look like?

1280x720 is a valid custom size for GPT Image 2.5 on Sume: both edges are multiples of 16, the ratio is 16:9 and the pixel count is above the 655,360 minimum. With two references, set the size explicitly instead of relying on auto. Set n above 1 only if you want to compare options; each image is billed.

curl -X POST "https://api.sume.com/v1/images" \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "openai/gpt-image-2.5",
    "prompt": "Image 1 is a layout sketch: use it only for composition, keep the positions and sizes of its boxes. Image 2 is the person: keep her face and identity exactly. Render a photorealistic YouTube thumbnail, bold headline \"I quit my job\" on the right, warm rim light.",
    "image_size": "1280x720",
    "quality": "high",
    "input_references": [
      {"type": "image_url", "image_url": {"url": "https://example.com/sketch.png"}},
      {"type": "image_url", "image_url": {"url": "https://example.com/portrait.jpg"}}
    ]
  }'

What should I check in the result?

Compare the output with the sketch for box positions and with the portrait for the face. Read the headline letter by letter. If the layout drifted, say so in the prompt (for example, name the exact side the text sits on) and shorten the style sentence. If the face drifted, restate the identity constraint, as covered in the same-face prompt block.

Sume does not guarantee that a sketch is followed pixel for pixel, and the platform you upload to sets its own thumbnail size and file limits; check those before exporting. For choosing among several options in one call, see AI thumbnail variants in one call.

How should I draw the sketch?

Keep it simple: boxes for the face and the text, a rough arrow or object shape, and nothing else. Label each box in the prompt, for example "left box is the face, right box is the headline". Dark lines on a white page are easier for the model to read than a photograph of a notebook.

Make the sketch the same shape as the output, 16:9 here, so the boxes do not need to be reinterpreted. If the first result misses the layout, fix the sketch before you rewrite the prompt.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume