Pick the best of 4 AI image takes automatically with Pillow

Request n=4 from one Sume image call, score each take by edge contrast with Pillow ImageStat, keep the winner, and see what the four takes cost.

5 min readSume
All posts

Ask for four takes in one request with n=4, score each with a sharpness number you compute in Pillow, and keep the highest. On Seedream 4.5 that is four images for $0.20, because the catalog list of $0.04 per image times Sume's 1.25 billing ratio is $0.05 each. The score is a cheap first filter, not a taste judge, so treat it as a way to drop blurry takes before a person looks.

Why filter takes in code at all

A single prompt returns takes that differ in focus, in how busy the background is, and in whether the product edge is crisp. If you generate a catalog of 200 products, nobody wants to open 800 files. A scoring pass cuts the pile to one candidate per product, and the person only reviews winners.

Sume's image route accepts n up to a per-model cap (4 on Seedream 4.5), so all four takes come from one call and one usage.cost figure. Models with a cap of 1 need four calls instead, and the code below would loop over calls.

The scoring function

Convert the take to grayscale, run ImageFilter.FIND_EDGES, and read the standard deviation of the result with ImageStat.Stat. A soft, out-of-focus image has weak edges, so its deviation is low. A crisp product on a clean background has strong, sparse edges and scores higher.

import os, io, requests
from PIL import Image
H = {"Authorization": "Bearer " + os.environ["SUME_API_KEY"]}
def gen(**body):
    r = requests.post("https://api.sume.com/v1/images", json=body, headers=H, timeout=60)
    r.raise_for_status()
    if r.status_code == 202:
        raise SystemExit("queued, read /v1/jobs/{id}/result: " + r.text)
    return r.json()
def fetch(u):
    return Image.open(io.BytesIO(requests.get(u, timeout=60).content))
from PIL import ImageFilter, ImageStat
out = gen(model="bytedance-seed/seedream-4.5", n=4,
          prompt="Matte ceramic mug on a plain linen table, soft window light")
def sharpness(img):
    edges = img.convert("L").filter(ImageFilter.FIND_EDGES)
    return ImageStat.Stat(edges).stddev[0]
takes = [fetch(d["url"]) for d in out["data"]]
scores = [sharpness(t) for t in takes]
best = scores.index(max(scores))
takes[best].save("best.png")
print("scores", [round(s, 1) for s in scores], "kept", best, "cost", out["usage"]["cost"])

What the four takes cost

The figures below come from the repo catalog (default tier, 1K output). Always log usage.cost from the response, which is the billed amount; the arithmetic here is only for planning.

List price times 1.25 for four takes, catalog values read 2026-10-05
Model idList per imageBilled per imagen=4 call
bytedance-seed/seedream-4.5$0.04$0.05$0.20
black-forest-labs/flux.2-pro$0.03$0.0375$0.15
google/nano-banana-2$0.08$0.10$0.40

Where an edge score misleads

  • A busy background (wood grain, fabric weave, foliage) scores high even when the product is soft. Put the product on a plain surface in the prompt, or score only a center crop.
  • Film grain and JPEG artifacts add edges. Compare takes at the same size and the same format.
  • The score cannot see a wrong logo, an extra handle, or bad hands. Keep a human on the final pick.
  • Ties are common on very clean takes. Break them by file size or by a second metric.

Make it part of a pipeline

Run the score on a 512 pixel center crop for speed, store the four scores next to the winner, and keep the losers for a week in case the reviewer disagrees. Pair it with the duplicate finder so you do not keep four near-identical winners across a batch, and cap spend with the budget guard.

If a call returns 202 because the 30 second wait budget ran out, read the images from GET /v1/jobs/{id}/result instead. The sample exits on that case so it never scores a half-finished job.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume