Do reference images cost extra on Sume? Flat rows vs token billing
On most Sume image rows, a reference photo adds nothing: the price is one flat per-image line. ChatGPT Image 2.5 bills tokens, so references add input cost.

Answer
On most Sume image rows the price is one flat output_image line per image, so a request with ten reference photos costs the same as a request with none. ChatGPT Image 2.5 is the exception: it is billed from output tokens plus estimated input tokens, so each reference raises the estimate a little.
You can see the flat line yourself. GET /v1/images/models/{id}/endpoints returns a pricing array with one entry, billable: output_image, unit: image, and a cost_usd that already includes the Sume margin.
| Row | Billing basis | Reference images change the price? |
|---|---|---|
| Seedream 4.0, 4.5, 5.0 Lite | flat per image | no |
| Flux 2 Pro and Flex | flat per image | no |
| Grok Imagine, Qwen Image | flat per image | no |
| Nano Banana 2 and Pro | flat per image, scaled by resolution | no, but resolution does |
| Ideogram V3 | flat per image | no |
| Ideogram 4.5 | flat per quality tier | no, but quality does |
| ChatGPT Image 2.5 | output tokens plus estimated input tokens | yes, a small amount |
Why it matters
- Budgeting a 1,000-image edit run on a flat row is simple: price x 1,000. On 2.5 you need to estimate quality, size and the number of references.
- Each row's reference ceiling is separate from the price. Grok Imagine and Qwen Image take 10, Ideogram 4.5 takes 5 and ChatGPT Image 2.5 takes 16.
- Text-only rows (Imagen 4, Recraft V4, Qwen Image Max, Soul) take zero and reject references.
Check it yourself
Run one request without references and one with, then compare usage.cost in each response. On a flat row, the two numbers match. On 2.5, the second can differ because input tokens are part of the estimate.
Sume rounds billable amounts up in USD micros, so the response cost may be a fraction of a cent above a hand calculation. Failed or cancelled generations are not billed at all. Docs: Sume Image API.
Before a large run
Prices and descriptors change when the catalog changes, so confirm them before you spend. Call GET /v1/images/models/{id}/endpoints for the row you plan to use and read its pricing line and supported_parameters; both come back in one response.
Then run a pilot of three to five images and read usage.cost on each response. Multiply by your planned count for a forecast you can trust. Completed generations are billed in full and failed or cancelled ones are not, so a pilot that errors costs nothing.
For big batches, use mode: "async" or mode: "webhook" with a public HTTPS webhook_url, so no request waits on the 30-second sync limit. Poll GET /v1/jobs/{id}/status and fetch GET /v1/jobs/{id}/result when the job completes.
Sources
Related posts
More in Pricing
- Draft at 480p, finish at 1080p: ten drafts and a final on Sume
Ten 5-second Wan 3.0 drafts at 480p plus one 1080p final cost $4.38 on Sume, against $13.75 for eleven 1080p runs. The arithmetic and its limit.
- Dubbing a 10-minute video: ElevenLabs Dubbing cost versus STT plus TTS
ElevenLabs lists Dubbing at $0.33 or $0.50 per minute (v1) and $2.20 (v2), so ten minutes is $3.30 to $22. A DIY chain of STT and TTS is about $0.53.
- Cheapest image edit on Sume: 14 edit-capable models ranked by price
From Grok Imagine at about 2.5 cents to ChatGPT Image 2 at about 26: all 14 edit-capable Sume image models, sorted by catalog list times 1.25.
- Eleven v4 free plan 10,000 credits vs Sume TTS at $0.0475 per 1,000
ElevenLabs lists a free plan with 10,000 credits a month. Sume TTS is pay as you go at $0.0475 per 1,000 characters. What 10,000 characters costs.
Written by Sume