GPT Image 2.5 on Sume: no quality means high, auto reserves max
Skipping quality on GPT Image 2.5 means high ($0.0355 at 1280x720). Sending auto reserves max ($0.1420). What that does to a budget.

On Sume, leaving quality out of a GPT Image 2.5 request means high. Sending quality: auto is different: the docs say auto reserves max for the estimate. At 1280x720 that is the gap between about $0.0355 and about $0.1420 per image, a factor of 4.
The Sume docs (read 2026-10-06) state both rules: if you omit quality the default is high, and auto quality reserves max.
Three ways to say 'no preference'
| You send | Treated as | Sume estimate |
|---|---|---|
| nothing | high | $0.0355 |
| quality: high | high | $0.0355 |
| quality: auto | reserve max | $0.1420 |
| quality: medium | medium | $0.0092 |
Why it matters
A reserve is an amount held before the job runs, so a batch of auto requests ties up more of your balance than it will necessarily spend. It also makes a pre-run estimate look much higher than the image cost you expected.
Auto also removes the one thing you control. Pin the quality you want. Low and medium are enough for drafts, and high is the finish. The xhigh and max post covers when the top tiers pay off.
Check it on your own account
Do not budget from a blog table alone. GET /v1/images/models lists every model with its descriptors, and GET /v1/images/models/{id}/endpoints shows the pricing line for one model. Then run one small request and read usage.cost on the response, which is the billed amount in USD; the token counts in usage are reported as 0 on this route.
Run the test at the quality and size you plan to ship, because both move the price. A single test at low quality costs under a cent for most sizes here, so it is a cheap way to confirm your assumptions before a batch.
Sync, async and failures
The /v1/images route waits up to 30 seconds for the image. If the job finishes in that window you get the result directly; otherwise you get a 202 and an async job to poll. Write your client to branch on the status code, since larger sizes and higher quality are the likely cases for a 202.
Requests are strict. A parameter the chosen model does not list returns 400 unsupported_parameter, stream returns a 400, and provider.only or provider.order accept only sume. Treat a 400 as a bug in the request, not a transient error, and do not retry it unchanged.
Unknowns
The docs describe the reserve, not the final billed amount for an auto job. Check usage.cost on a small run before you rely on a number.
Sources
Related posts
More in Developers
- Sume MCP consent: granting write always includes read
Sume's MCP consent offers mcp:read (required) and mcp:write (opt-in). A write grant always includes read, and there is no paid scope.
- Haystack MCPTool 30-second timeout vs Sume jobs_wait holds
Haystack's MCPTool times out tool calls at 30 seconds; Sume's jobs_wait holds 50 to 55. Raise invocation_timeout or wait in shorter slices.
- Heroku H12 at 30 seconds: call Sume async, not sync
Heroku's router ends a request at 30 s (H12). Sume sync mode waits up to 30 s too. Submit async, return 202 to the browser, then poll or take a webhook.
- Does Sume hosted MCP match the HTTP API? Where parity stops
Not fully. Hosted MCP covers most generation families, but Image 1.0 and Video 1.0 are REST-only, and schedules have no MCP tool. Use the HTTP API for those.
Written by Sume