Four 16:9 thumbnail options in one call: n=4 cost, and why n=5 fails
Set n to 4 for four 1536x864 options in one GPT Image 2.5 call: about $0.0420 at medium on Sume. n=5 returns a 400, because every model caps at 4.
Send n: 4 and one /v1/images call returns four options. At 1536x864 on GPT Image 2.5 that is about $0.0420 at medium and $0.1617 at high on Sume. n: 5 fails: every row in Sume's image catalog caps n at 4.
The error is a 400 invalid_request that reads "n must be between 1 and 4" followed by the public model id, and its details include the field and the min and max.
Cost for four
Billing scales with the image count, so four images cost four times one image at the same settings.
| Quality | One image on Sume | n=4 on Sume |
|---|---|---|
| low | $0.0045 | $0.0180 |
| medium | $0.0105 | $0.0420 |
| high | $0.0404 | $0.1617 |
Fitting more than four
For eight options, send two calls of four. Use async mode and polling if you send them together: the default wait on /v1/images is 30 seconds, and slow configurations fall back to a 202 job instead of a 200 body. Branch on the status code.
A budget loop that stops when usage.cost passes a limit works well here; the four-in-one-call post has the quality-by-quality arithmetic at a square size.
When four is the wrong number
Four options at high quality cost more than one pick at high plus three drafts at low. If you only need a direction, draft at low and finish the winner. The thumbnail variants post shows the polish step.
Check it on your own account
Do not budget from a blog table alone. GET /v1/images/models lists every model with its descriptors, and GET /v1/images/models/{id}/endpoints shows the pricing line for one model. Then run one small request and read usage.cost on the response, which is the billed amount in USD; the token counts in usage are reported as 0 on this route.
Run the test at the quality and size you plan to ship, because both move the price. A single test at low quality costs under a cent for most sizes here, so it is a cheap way to confirm your assumptions before a batch.
Sync, async and failures
The /v1/images route waits up to 30 seconds for the image. If the job finishes in that window you get the result directly; otherwise you get a 202 and an async job to poll. Write your client to branch on the status code, since larger sizes and higher quality are the likely cases for a 202.
Requests are strict. A parameter the chosen model does not list returns 400 unsupported_parameter, stream returns a 400, and provider.only or provider.order accept only sume. Treat a 400 as a bug in the request, not a transient error, and do not retry it unchanged.
Sources
Related posts
More in Developers
- Gemini CLI and the hosted Sume server: do not rely on env in headers
Add the hosted Sume server to Gemini CLI with httpUrl and a bearer header. Gemini expands env vars only in the env block; set a timeout above jobs_wait.
- Gemini Omni 4K on Sume: send 4K or 4k, 10 s max, 16:9 or 9:16
Sume's gemini-omni-flash-1.1 takes 4K, and lowercase 4k works as an alias. Requests run 3 to 10 seconds, 16:9 or 9:16. Python body builder.
- GET /v1/video-router/models/wan-3.0: fetch one model's limits
Fetch one Sume video model with GET /v1/video-router/models/{model_id}, such as wan-3.0 or seedance-2.5. Unknown ids return 404. Read limits before you submit.
- Get the video URL from a Sume webhook: pick the artifact by type
A Sume job.completed payload lists artifacts with id, url, type and content_type. Select the video by content_type, not array index. TypeScript for Node.
Written by Sume