24 character stills on gpt-image-2.5: 24c low, 48c medium, $1.68 high
If you omit quality, gpt-image-2.5 defaults to high. 24 stills (8 characters, 3 poses) cost 24c at low, 48c at medium and 144 to 168c at high on Sume.

Twenty-four character stills on gpt-image-2.5 cost 24 cents at low quality, 48 cents at medium and between 144 and 168 cents at high, using the 1024-class billable prices of 1c, 2c and 6 to 7c per image. The catch is the default. The Sume docs say that if you omit quality, it is high, so a request without the field is billed at the highest of the three.
The arithmetic
The set is 8 characters with 3 poses each, which is 24 stills. The per-image figures are the billable prices at 1024-class sizes.
| Quality | Per image | 24 images |
|---|---|---|
| low | 1c | 24c |
| medium | 2c | 48c |
| high (the default if you omit it) | 6c to 7c | 144c to 168c = $1.44 to $1.68 |
Use low to find the pose, high to keep it
A character sheet has two jobs. First you need to find the face and outfit that work. For that, low and medium are enough: 24 low stills cost 24 cents, so you can try several looks. Then you keep the winners and render them at higher quality. The model supports low, medium, high, xhigh and max, but the docs price xhigh and max higher: at 1024x1024 their output alone is $0.09366 and $0.21072 before input tokens and Sume pricing, so they are not part of this table.
Each request returns up to 4 images, and the model takes up to 16 reference images. That second number is the useful one for consistency. Send the chosen face as a reference with each new pose, and the 16 slots hold the face, the outfit and the setting for the same character.
- Draft all poses at
quality: low. - Pick one still per character that looks right.
- Re-render the keepers at
high, with the draft as a reference image. - Reuse those stills as references in video clips.
Using the stills afterwards
Once you have 8 approved stills, they become reference images for video. The Sume video catalog lists reference images on Wan 3.0 (up to 10), MiniMax H3 and H3 Max (up to 9) and Gemini Omni Flash 1.1 (up to 10). A character with a front still and a three-quarter still needs 2 of those slots. That leaves room for a second character and a setting still in the same clip.
Keep the file names stable. A still called character-3-front is easier to find in a list of reference URLs than a job id. All reference URLs must be public HTTPS; localhost and private-network URLs are rejected before submission.
When high is worth the extra
Higher quality earns its price when text or fine detail must read: a logo on a jacket, a name on a badge, small print on a prop. For a face and body that will be a reference, medium is often enough, since the video model reads the overall look. Decide once per character, and do not let the default decide for you. At 168 cents for 24 stills the cost is small, but at 1,000 stills it is up to $70 at high, $20 at medium and $10 at low.
Always set the field
Setting quality is one line of the request. The code builds the requests for the 8 x 3 grid and refuses to build one without it, so that nobody ships a default by accident.
import json
CHARACTERS = [f"character-{i}" for i in range(1, 9)]
POSES = ["front view", "three-quarter view", "walking"]
CENTS = {"low": 1, "medium": 2}
def request(name: str, pose: str, quality: str) -> dict:
if quality not in CENTS:
raise ValueError("pick low or medium for drafts")
return {"model": "openai/gpt-image-2.5", "quality": quality,
"prompt": f"{name}, {pose}, plain studio background"}
reqs = [request(c, p, "low") for c in CHARACTERS for p in POSES]
print(len(reqs), "stills, cost", len(reqs) * CENTS["low"], "cents")
print(json.dumps(reqs[0]))Sources
Related posts
More in Media tools
- A 30-second music bed under a 70-second render costs $0.325
One Music Router track at $0.125 plus a 70-second Timeline render at $0.20 is $0.325. Use soundtrack loop, fade_out_seconds up to 10 and duck_db 0 to 20.
- 45-minute recording: audio detach fails past 1800 s, what to do
Sume audio detach accepts a source up to 1800 seconds and returns up to 900. A 45-minute video is 2700 s, so trim first. Here is the arithmetic.
- 45-second voice-over: four H3 Max segments, $4.80, or one Fabric job
A 45-second line must be four H3 Max lip-sync segments (about $4.80 at 768p) or one Fabric 1.0 job (about $8.44 at 720p). What each route allows and costs.
- 90 s Audio Library song in a 3-minute Short: what covers the rest
YouTube lets a Shorts Audio Library song play up to 90 s in a 3-minute Short. Sume Music 1.0 is $0.125 a track; Timeline loops a bed for $0.30.
Written by Sume