Logo sketch to a transparent PNG with GPT Image 2.5
Turn a hand-drawn logo sketch into a transparent PNG: send the sketch as a reference, set background transparent and output_format png, then verify alpha.

To get a transparent logo from a sketch, send the sketch as an input_references entry to GPT Image 2.5, set background to transparent and output_format to png, then check the alpha channel of the file you get back. On Sume only the two GPT Image 2.5 variants accept background.
Treat the result as a draft mark, not a finished brand. Raster output has soft edges, so a logo that must scale to a billboard still needs a vector redraw.
Why png or webp
The OpenAI image guide says transparent backgrounds need an output format that supports alpha, which means png or webp. JPEG has no alpha channel. So request png for a logo you will place on other colors.
Describe the mark flat. Ask for solid shapes, two or three colors, no gradient, no shadow and no mockup scene. A shadow turns into semi-transparent pixels that look dirty on a dark background.
| Field | Value | Note |
|---|---|---|
| model | openai/gpt-image-2.5 | Variants only: background is off elsewhere |
| background | transparent | Other models answer 400 unsupported_parameter |
| output_format | png | Alpha needs png or webp |
| input_references | 1 public HTTPS sketch | Pencil or marker on white |
| image_size | 1024x1024 | Both edges multiples of 16 |
Request and alpha check
Reference URLs must be public HTTPS, and Sume answers 400 input_media_unreachable when it cannot download one. A file on your laptop needs a public upload first, for example through the assets upload flow.
import os, requests
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
def generate(body):
r = requests.post("https://api.sume.com/v1/images", headers=H, json=body, timeout=60)
if r.status_code != 200: # 202 = still running, read data.status_url
raise SystemExit(f"{r.status_code}: {r.text[:300]}")
return r.json()["data"][0]["url"]
from io import BytesIO
from PIL import Image
url = generate({
"model": "openai/gpt-image-2.5",
"prompt": "Clean flat logo from this sketch, two colors, no shadow, no gradient",
"input_references": [{"type": "image_url", "image_url": {"url": "https://example.com/logo-sketch.jpg"}}],
"background": "transparent",
"output_format": "png",
"image_size": "1024x1024",
})
im = Image.open(BytesIO(requests.get(url, timeout=60).content))
low = im.convert("RGBA").getchannel("A").getextrema()[0]
print("has transparency" if low < 255 else "opaque: check the prompt")
im.save("logo.png")Price
GPT Image 2.5 is billed on tokens. The fal pages list $30 per million output image tokens, $8 per million image input tokens and $5 per million text input tokens. Sume bills the provider list price times 1.25. The table is output-only, so reference and prompt tokens add a little on top of it.
| Quality | Provider list | Sume at list x 1.25 |
|---|---|---|
| medium | $0.0132 | $0.0165 |
| high | $0.0527 | $0.0658 |
Check on dark and light
Paste the PNG on a black and a white square before you trust it. Fringing, a faint box around the mark or leftover paper texture shows up on one of the two.
If the call returns 202
POST /v1/images waits up to 30 seconds and returns 200 with the images. A slow job falls back to a 202 job envelope, and you read the images from GET /v1/jobs/{id}/result. The code above exits on any non-200 so you notice, and a failed synchronous job returns 502 and is not billed.
When the file is not transparent
Three causes explain most opaque results. A short checklist finds them quickly. If the three checks pass and the file is still opaque, read the response media_type and re-run with png, then open the file in an editor that shows a checkerboard behind transparent pixels.
- Confirm the request carried
background: "transparent"and anoutput_formatofpngorwebp. - Remove words such as "on a white background" or "mockup" from the prompt.
- Check the model id. Off GPT Image 2.5,
backgroundanswers400 unsupported_parameter.
Sources
Related posts
More in Use cases
- Voice drift in long narration: Nova 2 Sonic's 52% vs Sume chunks
Speaker drift is what splits a long voiceover into jobs. Amazon claims -52% on an internal set; here is how to measure drift across Sume TTS chunks yourself.
- Luggage holiday travel ad: fit a 16:9 clip to 9:16 with blur for $0.11
Reuse a 16:9 luggage ad as a 9:16 Reel: detach the audio for $0.01, then one Timeline render with fit blur for $0.10, instead of $1.00 to regenerate.
- Live transcription for Korean meetings: MAI-Transcribe-2-Streaming
MAI-Transcribe-2-Streaming claims 60 languages with auto detection and about 100ms to first text. Verify Korean yourself. Sume transcribes files, not live.
- Make an AI voice read an email or order code right: spell, verify
Amazon and Cartesia both claim better codes, emails and phone numbers. The reliable fix is text prep plus a read-back check. Try both on Sume's TTS router.
Written by Sume