Does a longer prompt cost more on GPT Image 2.5? Text token math
Prompt text is billed at $5 per million tokens on GPT Image 2.5, so a 2,000-token brief adds about $0.01, as much as a medium image. Numbers and rules.

Yes, on GPT Image 2.5 a longer prompt costs more, but not by much per word. Fal lists text input at $5 per million tokens, so a 2,000-token brief adds about $0.01 to an image. That is about the same as the whole output estimate for a medium 1024x1024 image ($0.01317). For a short prompt the text is a rounding error; for a long style brief sent on every call, it can double a cheap draft.
The arithmetic
Fal's GPT Image 2.5 Flare page lists text tokens at $5.00 per million input, $1.25 cached, $10.00 output, and image tokens at $8.00 input, $2.00 cached, $30.00 output. It also says longer prompts increase the cost. The Sume docs use the same Fal token rates and say input token counts are estimates. Token counts for English run roughly 1.3 tokens per word, which is our approximation, not a figure from either vendor.
| Prompt | Approx. tokens | Text cost per image |
|---|---|---|
| 20 words | 26 | $0.00013 |
| 150 words | 195 | $0.00098 |
| 500 words | 650 | $0.00325 |
| 1,500 words | 1,950 | $0.00975 |
Where it starts to matter
Compare against output: at 1024x1024 the output estimate is $0.00588 for low, $0.01317 for medium, $0.05268 for high. A 500-word brief adds $0.00325 which is half of a low image and a quarter of a medium one. On high or above it vanishes.
The pattern is the one that bites: a shared 500-word brand brief pasted into 1,000 draft calls adds $3.25. Moving the stable part of the brief out of the prompt and into a reference image is not free either, because input image tokens are billed at $8 per million.
Ways to keep it small
- Put the subject first and the style rules after it, and cut anything the model already does by default.
- Reuse one reference image for style instead of a paragraph describing it; read the estimate for the reference tokens on a test call.
- Draft at
low, and write the long brief only for the final pass. - Keep a house style as a short string in your wrapper, not a pasted document.
Verify on your own account
Cached input is cheaper on Fal's list ($1.25 text, $2.00 image), but whether Sume applies a cached rate to your repeated prompts is not something the Image API docs promise, so do not budget on it. Send the same call with a short and a long prompt at n: 1 and subtract the two usage.cost values. That difference is your real price per extra word on Sume.
Sources
Related posts
More in Pricing
- GPT Image 2.5 quality auto reserves max on Sume: pin quality
Sume's image docs say quality auto reserves max, and auto size reserves the output token upper bound. What to set instead, plus the 1024 by 1024 figures.
- Grok Imagine: 20 MiB image cap and per-second prices
xAI lists Grok Imagine video at $0.020 to $0.080 a second and a 20 MiB image limit. On Sume, check the live catalog and add the x1.25 billing rule.
- Grok Imagine Image 2.0: auto means low, then medium
xAI's Grok Imagine Image 2.0 price list by quality, what auto resolves to, and how to read a model's price lines on Sume before you spend.
- Grok Imagine Video 1.5: cost of 10 seconds
At xAI's listed per-second rates, a 10-second Grok Imagine Video 1.5 clip costs $0.80 at 480p, $1.40 at 720p and $2.50 at 1080p. Check Sume's catalog prices.
Written by Sume