Near-duplicate AI images: perceptual-hash dedupe before human review

Four-up image batches often contain near twins. Use a 64-bit difference hash in Pillow to drop duplicates before a person reviews them. Python for Sume results.

5 min readSume
All posts

To remove near-duplicate AI images before review, compute a 64-bit difference hash (dHash) for each result and drop any image whose hash is within a few bits of one you kept. A dHash shrinks the picture to 9 by 8 grayscale pixels and records whether each pixel is brighter than its right neighbor. Two renders of the same scene produce hashes a few bits apart, so a Hamming distance of about 6 or less is a good first threshold to tune on your own batch.

This matters when you ask Sume for several images per call. The n parameter takes 1 to 10 at the API level, with a lower ceiling on some models, and a batch from one prompt often returns variations that look the same to a reviewer who is paid by the hour.

How does the hash compare two images?

Distance guide for a 64-bit dHash, read 2026-10-04
Hamming distanceTypical meaningAction
0 to 3Same image, maybe re-encodedDrop
4 to 6Same composition, small differencesDrop or keep one
7 to 12Related but distinctKeep both
13 and upDifferent imageKeep

What is the code?

Feed it a list of image URLs from data[].url. It returns the URLs to keep. Pillow is the only dependency.

import io, requests
from PIL import Image

def dhash(img):
    g = img.convert("L").resize((9, 8), Image.LANCZOS)
    px = list(g.getdata())
    bits = 0
    for row in range(8):
        for col in range(8):
            left, right = px[row * 9 + col], px[row * 9 + col + 1]
            bits = (bits << 1) | (left > right)
    return bits

def unique(urls, limit=6):
    kept = []
    for u in urls:
        img = Image.open(io.BytesIO(requests.get(u, timeout=60).content))
        h = dhash(img)
        if all(bin(h ^ k).count("1") > limit for k, _ in kept):
            kept.append((h, u))
    return [u for _, u in kept]

if __name__ == "__main__":
    print(unique(["https://example.com/a.png", "https://example.com/b.png"]))

What does this save?

It does not reduce the generation bill, since Sume bills each completed image. What it saves is review time. If a four-image call returns two usable variations, a reviewer sees two, not four, and your approval-rate numbers become honest. See cost per approved ad image for the cost side.

Where does a hash fall short?

A hash sees layout and brightness, not meaning. Two images of the same product with different backgrounds can hash far apart, and two different products on the same plain background can hash close. Use it to cut obvious twins, not to pick a winner. For a catalog, compare only within one prompt's batch, never across unrelated products.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume