Which image model follows a rough sketch best? A test plan on Sume
No honest ranking exists without your sketches. Run one sketch and one prompt through four Sume image models, score six traits, and read the cost per try.

There is no public ranking that tells you which image model follows a rough sketch best for your drawings, so run the test yourself: one sketch, one prompt, four models, six scored traits. On Sume, the sketch goes in as an input reference; there is no separate sketch mode. GPT Image 2.5 takes up to 16 references and a mask, and OpenAI's guide describes the model family, but how well each model keeps your composition depends on the drawing (OpenAI).
Set up
Pick one clear sketch and one prompt that names the layout role of the reference: "Image 1 is a rough sketch. Keep the composition. Render it as a photo." Send the same request to each model with aspect_ratio: "auto" where listed, and run each twice, because an image model gives different results per call and Sume has no seed field (Image API).
| Model id | Sume rate per image | Notes |
|---|---|---|
| openai/gpt-image-2.5 | Token based | Up to 16 references, mask_url |
| google/nano-banana-pro | $0.19 at 1K | Lists aspect_ratio auto |
| black-forest-labs/flux.2-pro | $0.04 | Edits from references |
| bytedance-seed/seedream-5-lite | $0.05 | Edits from references |
Score six traits
Give each result 0 to 2 on each trait and total them. Do it blind if a colleague can shuffle the file names.
- Composition kept: are the main shapes where the sketch put them?
- Proportions kept: sizes and angles of objects.
- Count kept: the number of objects, people or windows.
- Text and labels: do written words match your quoted text?
- Style fit: does it look like the style in the prompt?
- Surprises: added or missing elements.
Read the result
Add up the totals, then divide cost by usable results rather than by tries. A model that is half the price but needs three attempts is not cheaper. If the winner differs between sketches, route by sketch type: loose scribbles to one model, tidy line drawings to another. Check the live price with GET /v1/images/models/{id}/endpoints.
Sources
Related posts
More in Comparisons
- Wix Stores promo video maker vs a custom AI product clip
Wix Stores can auto-make a promo video per product, free. When is a custom Sume image-to-video clip from your own photo worth the extra step? A side-by-side.
- YouTube reused vs inauthentic content: which rule hits AI Shorts?
Two YouTube monetization rules and one new Shorts reach change, side by side. Which one an AI-generated Short can trip, and what to vary to avoid each.
- Sume vs Argil: AI avatar video and video agents compared
Argil makes AI-avatar and story videos with a chat agent, Director; Sume is a video agent with a multi-model API. Avatars, API, pricing, and limits compared.
- Sume vs fal: a generative media API or a video agent platform
fal runs 1,000+ image, video, and audio models behind one API. Sume adds a video agent, Formats, and avatars to a multi-model API. How the two surfaces differ.
Written by Sume