Shrink an AI image to a byte budget: JPEG quality search in Pillow
Platforms cap image files at 1 MB or 5 MB. Binary-search the JPEG quality in Pillow for the sharpest file under your byte limit, from a Sume image.

Save the image to an in-memory buffer at a trial quality, read the buffer length, and halve the search range between 30 and 95 until the largest quality under your byte limit is found. Seven saves are enough, because the range is 65 wide and each step halves it. The loop runs in memory, so it does not touch disk until the winner is written.
Why a fixed quality is a bad answer
The same quality number gives very different file sizes. A smooth gradient at quality 85 can be 80 KB, and a detailed texture at quality 85 can be 900 KB. If a form accepts at most 1 MB, a fixed number either wastes quality or fails the upload.
A search finds the best quality for this image. Since a Sume image result is a PNG or another format at the model's own size, you also decide the pixel size. Resize first, then search quality.
The search
Pillow's JPEG writer takes quality from 1 to 95 (values above 95 are discouraged), plus optimize=True for smaller files and progressive=True to render coarse-to-fine. File size rises with quality, which is what makes binary search valid.
import os, io, requests
from PIL import Image
H = {"Authorization": "Bearer " + os.environ["SUME_API_KEY"]}
def gen(**body):
r = requests.post("https://api.sume.com/v1/images", json=body, headers=H, timeout=60)
r.raise_for_status()
if r.status_code == 202:
raise SystemExit("queued, read /v1/jobs/{id}/result: " + r.text)
return r.json()
def fetch(u):
return Image.open(io.BytesIO(requests.get(u, timeout=60).content))
out = gen(model="bytedance-seed/seedream-4.5", aspect_ratio="4:3",
prompt="Detailed cobblestone street in a European old town, afternoon light")
img = fetch(out["data"][0]["url"]).convert("RGB")
LIMIT = 300_000
def size_at(q):
buf = io.BytesIO()
img.save(buf, "JPEG", quality=q, optimize=True, progressive=True)
return buf.tell(), buf.getvalue()
lo, hi, best = 30, 95, None
while lo <= hi:
mid = (lo + hi) // 2
n, data = size_at(mid)
best, lo, hi = ((mid, n, data), mid + 1, hi) if n <= LIMIT else (best, lo, mid - 1)
if best is None:
raise SystemExit("even quality 30 is over the limit; shrink the pixels first")
open("out.jpg", "wb").write(best[2])
print("quality", best[0], "bytes", best[1], "cost", out["usage"]["cost"])Typical limits
Limits come from each platform's own upload rules, so check the current page before you rely on a number. The ones below are the usual shape of the problem.
| Quality range | Width | Saves needed (ceil log2) |
|---|---|---|
| 30 to 95 | 66 values | 7 |
| 1 to 95 | 95 values | 7 |
| 60 to 90 | 31 values | 5 |
When quality alone is not enough
- If quality 30 is still over the limit, lower the pixel size and search again.
- Chroma subsampling of 4:2:0 is the default for lower qualities and saves bytes on photos. Fine red text can smear.
- Avoid re-saving the JPEG. Always compress from the lossless source, as in PNG between passes.
- Keep the PNG master for later edits.
Put it in the pipeline
Run the search as the last step after cropping to the target ratio, for example the 1080x1350 feed image, and generate other widths from the same master with the srcset recipe. Read usage.cost from the same response to record what each final file cost.
Sources
Related posts
More in Developers
- silence_split_seconds: tune caption line breaks from STT segments
Sume STT sentence segmentation can split on silence. silence_split_seconds takes 0.2 to 3 and returns gapless segments you can use as caption lines.
- Captions fail on a silent Short with caption_no_speech: send cues
A silent clip has no speech to transcribe, so Sume's captions API returns caption_no_speech. Send cues with text, start and end to burn overlay text instead.
- 16 AI shots in one Sume Timeline render: the 8-fade cap
A 16-shot cut fits one Timeline render, but fades are capped at 8 in a row and renders chunk past 12 slots. Plan the cuts, with the doc limits.
- Size a batch from generation_limits so no clip hits queue_full
Read accepted_generation_jobs_limit from a Sume submit response and slice your clips. A 50-clip batch leaves 2 for a later wave on Startup and 26 on Pro.
Written by Sume