A ten-photo edit regression suite to re-run when a model launches
Ten photos, five edits each, one scoring sheet: a cheap test to re-run whenever a new image edit model launches. 50 edits cost $1.875 at the low tier on Sume.

Keep a fixed set of ten photos and five edit prompts, and run it on each new model before you believe its launch page. Fifty edits on Ideogram 4.5 at the low tier cost $1.875 on Sume ($0.0375 each), and $3.75 at medium. Flux 3 Image and Ideogram 4.5 both arrived within a day of each other with claims about unchanged pixels, and the only way to compare them on your own work is a repeatable test.
The ten photos
Choose photos where drift is easy to see: faces, small text and fine textures. Store them as PNG in a folder under version control, with a file for each prompt.
- A product on white, a product on a busy table, a person facing the camera, a group shot.
- A room interior with a window, a street scene with signs, a food shot with steam.
- A flat graphic with text, a packaging photo with small print, a very large photo.
The five edits
Run each edit alone, then chain three on one photo to look for drift. Record the model id, the quality tier, the parameters and the output URL for every run.
- Recolor one object and keep another unchanged.
- Change a line of text and keep the lettering.
- Remove one object and keep its surroundings.
- Relight to a different time of day.
- Add an object at a named place.
Score it
Use four yes or no checks per edit: the change is right, the named keep is intact, the rest of the frame is unchanged (diff it), and the text reads correctly where text is involved. Fifty edits give you up to 200 ticks.
| Quality | Per image | 50 edits | Add a 3-pass chain on all 10 photos (30 more) |
|---|---|---|---|
| low | $0.0375 | $1.875 | $3.00 |
| medium | $0.075 | $3.75 | $6.00 |
| high | $0.275 | $13.75 | $22.00 |
What to compare against
Run every model with the same seed-free settings and the same quality tier, or note where they differ, because a high tier against a low tier is not a comparison of models. Keep the first result you get for each prompt; picking the best of three on one model and the first on another tilts the table.
A launch page shows its best results, with the unchanged shares BFL chose to publish for Flux 3 Image (cat removal 80.7%, bird addition 86.4%, shirt print 89.7%, hair recolor 67.8%, per TechTimes) and Ideogram's prose about untouched pixels. Your suite shows how a model behaves on your photos. Compute the unchanged share the same way for each model: the fraction of pixels outside your edit region that are identical, as in the pixel diff recipe.
Note which models the Sume catalog lists before you run. Flux 3 Image is not among them as of this read, so it cannot be run through the Sume Image API; test it on the vendor's own service if you need the comparison, and check the Sume list again after any launch.
For the catalog check, see a Python script for references and masks.
Sources
Related posts
More in Developers
- Test a Sume webhook receiver with signed fixtures, no paid job needed
Generate sume-v1 signatures yourself and test six cases: good, rotated, reserialized, stale, empty-secret and unknown event. Python code that runs as is.
- Thai, Vietnamese, Indonesian speech to text: Sume STT language hints
Send th, vi or id as language_code to Sume STT, or omit it to auto-detect, then check language_probability. $0.01 per audio minute; test a sample first.
- Timeline output.fps: why a 24 fps clip judders when you force 30
Leave output.fps unset and Timeline renders at the source rate. Force 30 on a 24 fps clip and frames repeat; the job reports output_fps_resamples_sources.
- Transcribe a file from your own bucket: what Sume STT audio_url needs
Sume speech-to-text takes a public HTTPS audio_url, preferably on media.sume.com. How to get a private recording ready, and what a neighbouring API rejects.
Written by Sume