Price a mixed image batch first: endpoint pricing in Python
Read cost_usd from each Sume image endpoint and multiply by the count. 200 Grok, 50 Seedream 4.5 and 100 Flux 2 Pro images should total $11.25 before you spend.

Sume publishes the per-image price of every image model at GET /v1/images/models/{id}/endpoints, in a pricing list of {billable, unit, cost_usd} lines. Multiply cost_usd by the number of images and you have the bill before you send a generation. For 200 Grok Imagine, 50 Seedream 4.5 and 100 Flux 2 Pro images the answer should be $11.25.
This is different from a budget guard, which stops a run after spending starts. An estimate tells you the size of the run first, so you can change the plan while it costs nothing.
What the endpoint returns
The Image API docs show an endpoint record with provider_name: "Sume" and a pricing array such as {"billable": "output_image", "unit": "image", "cost_usd": 0.033}. The OpenAPI description for that field says cost_usd is the amount charged to the caller, margin included. In other words, the number is already list x 1.25; do not multiply it again.
The docs also say that Sume serves every catalog model through a single sume endpoint, so endpoints[0] is the only record. The pricing list has one line for flat-price models. Models that price by tier or quality, such as Nano Banana and Ideogram 4.5, show the catalog default, so for those the estimate is a floor for the tier you request.
Expected numbers for the example plan
If the endpoints report the billed prices Sume documents for these models, the script should print the lines below. A difference means the catalog changed; trust the endpoint over this table.
| Model | Images | Billed per image | Subtotal |
|---|---|---|---|
| x-ai/grok-image | 200 | $0.025 | $5.00 |
| bytedance-seed/seedream-4.5 | 50 | $0.05 | $2.50 |
| black-forest-labs/flux.2-pro | 100 | $0.0375 | $3.75 |
| Total | 350 | $11.25 |
The script
It needs only the standard library and an API key in SUME_API_KEY. It asks each endpoint for its first pricing line and multiplies by the planned count.
import json, os, urllib.request
BASE = "https://api.sume.com/v1/images/models/"
KEY = os.environ["SUME_API_KEY"]
def unit_price(model):
req = urllib.request.Request(BASE + model + "/endpoints",
headers={"Authorization": "Bearer " + KEY})
with urllib.request.urlopen(req, timeout=30) as r:
info = json.load(r)
return float(info["endpoints"][0]["pricing"][0]["cost_usd"])
plan = {
"x-ai/grok-image": 200,
"bytedance-seed/seedream-4.5": 50,
"black-forest-labs/flux.2-pro": 100,
}
total = 0.0
for model, count in plan.items():
price = unit_price(model)
cost = price * count
total += cost
print(f"{model}: {count} x {price:.4f} = {cost:.2f}")
print(f"estimate: {total:.2f}")
Why this beats a spreadsheet
A spreadsheet of prices goes stale the day the catalog changes. The endpoint is the source the billing path reads, and its cost_usd field is documented as the amount charged, margin included. Reading it at run time means your estimate and your invoice come from the same number.
It also scales. Adding a fourth model to the plan is one more line in the dictionary, and removing one is deleting a line. The totals update on their own, so you can try three or four plans in a minute and pick the one that fits your budget before any generation starts.
Keep the estimate honest
- Run the script from the same account that will pay for the batch, since the endpoint answers for the key you send.
- Re-read prices at run time. Do not paste them into a spreadsheet and trust them next month.
- For tiered models, check the line for the tier you will request, or compare the estimate with
usage.coston a single test image. - Add a margin for retries only after you know your own failure rate. Failed generations are not billed.
- After the real run, compare the summed
usage.costwith the estimate and log any gap. - Keep the plan dictionary in version control next to the prompts, so a reviewer can see what the batch was meant to cost.
Sources
Related posts
More in Developers
- Fade in and out on a Timeline render: output fade seconds 0 to 5
Set output.fade_in_seconds and fade_out_seconds (0 to 5 s, sum within the render length). The music bed has its own fade_out_seconds, up to 10.
- Fast-cut Shorts in Timeline: eight chained fades, then a hard cut
Timeline refuses more than 8 adjacent fades with too_many_chained_transitions. Transitions must be 1 s or less and half the shorter neighbour. How to plan cuts.
- FastMCP 4 client OAuth with Sume: request mcp:read, add write later
Use FastMCP's OAuth helper in Python to sign in to Sume's hosted MCP with mcp:read first, then ask for mcp:write only for the script that spends.
- Fix an underexposed photo: curves first, AI edit only if needed
Underexposed photo? Try Pillow autocontrast and gamma for free, then an ideogram/ideogram-v4.5 edit at $0.075 only if noise or colour needs more.
Written by Sume