Virtual staging an empty room photo with AI: GPT Image 2.5 via API
Stage an empty room photo with openai/gpt-image-2.5: the room plus up to 15 furniture references, an optional mask, and about $0.08 per staged image at high.

To virtually stage an empty room with an API, send the room photo plus photos of the furniture you want to see in it to POST /v1/images with model: "openai/gpt-image-2.5", and describe the layout in the prompt. ChatGPT Image 2.5 on Sume takes up to 16 images in input_references, so one room and up to 15 furniture or style references fit in a single call. At quality: "high" a one-reference edit at 1024x1024 prices at $0.0835 in Sume's pricing fixtures, which is about eight cents for one staged image.
The rest of this post covers how to structure the request, when to add a mask, how the cost moves with more references, and a check that the walls and windows did not move.
What the request looks like
Staging is an edit with references, not a text-to-image call. The room is the thing that must not change, and the furniture photos are the things that must appear. The Sume docs say the reference limit on ChatGPT Image 2.5 is 16, that reference URLs must be public HTTPS, and that localhost, private-network and non-HTTPS URLs are rejected before submission.
State in the prompt which image is the room. The docs do not define an order rule for this model, so do not rely on one. Write something like: the first image is the room, keep walls, windows, floor and camera angle exactly as they are, add the sofa from image two and the rug from image three. Then check the result instead of trusting the sentence.
curl -X POST https://api.sume.com/v1/images \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "openai/gpt-image-2.5",
"prompt": "Image 1 is an empty living room. Keep walls, windows, floor and camera angle unchanged. Place the sofa from image 2 against the long wall and the rug from image 3 under it.",
"input_references": [
{"type": "image_url", "image_url": {"url": "https://example.com/empty-room.jpg"}},
{"type": "image_url", "image_url": {"url": "https://example.com/sofa.jpg"}},
{"type": "image_url", "image_url": {"url": "https://example.com/rug.jpg"}}
],
"aspect_ratio": "auto",
"quality": "high"
}'Mask or no mask
mask_url is live only on ChatGPT Image 2.5 in the Sume catalog, so staging is a place where a mask is available if you want it. A mask limits the change to the floor area where the furniture goes and leaves the ceiling and windows out of the model's reach. The docs describe mask_url only as a public HTTPS mask URL for edits; they do not spell out the colour convention, so build a small test mask, run one cheap low edit, and diff the output against the source before you batch.
Use aspect_ratio: "auto" on edits so the output follows the reference. The Sume image docs say that if you omit the field the result is not the same as auto. For a listing photo you almost always want the original frame.
Cost by quality, one reference
The pricing fixtures in packages/provider-pricing hold the billed amounts for a 1024x1024 edit with one input image. More references add input tokens, so treat these as a floor for a staging call that carries three or four furniture photos.
| Quality | Billed per image | 20 rooms, one style |
|---|---|---|
| low | $0.009375 | $0.1875 |
| medium | $0.021 | $0.42 |
| high | $0.0835 | $1.67 |
| xhigh | $0.148375 | $2.9675 |
Prompt habits that help
Name the room type and the camera position once. Give each furniture reference a role: sofa, rug, floor lamp. Tell the model the scale you expect, for example that the sofa is about two metres wide against a four-metre wall. Models invent scale when you leave it out, and an oversized sofa is the most common failure in staged photos.
Ask for lighting that matches the room: the same window light direction, soft shadows under the legs, no new light sources. Keep each call to one style. If you want three styles of the same room, make three calls rather than one call that asks for all three, since mixing styles in one prompt confuses the references.
Check that the room did not move
A staging edit is only useful if the architecture is untouched. Download the source and the result at the same size, take the absolute difference, and look at where the difference is. Furniture and its shadows should light up. Window frames, door positions and ceiling lines should not. If the window moved, rerun with a mask or move to a lower-change prompt; a completed edit is billed whether or not you like it, and failed jobs are not billed.
Staged images should also be labelled as virtually staged wherever the listing platform asks for it. That is a listing rule, not a model setting, so keep the original photo next to the staged one in your asset folder.
When to pick another route
If the furniture does not need to match real products, drop the references and describe the style; the call is cheaper because there are no input image tokens. If you need an exact, repeatable layout, an AI edit will not give you a seed, since seed is not served on Sume in v1 and returns 400 unsupported_parameter. Keep the best result and store its URL instead of trying to regenerate it.
Sources
Related posts
More in Use cases
- How do I make a winter tire changeover promo video for my garage?
Three 6-second Wan 3.0 clips, a 420-character voiceover, a duck-under music bed and booking captions for a garage: $2.695 on Sume at 720p.
- Write a lip-sync script in 5 to 14.8 second beats (H3 Max voice-over)
MiniMax H3 Max lip-sync on Sume takes audio of 5 to 14.8 seconds. Write the script in beats that fit the window, then voice each beat; the price per beat.
- YouTube bumper ad 5-6 seconds: generate at 6 or trim on Sume
YouTube's help page puts bumper ads at 5 to 6 seconds. Which Sume video models can generate a 6 second clip directly, and when a $0.02 trim is the better route.
- YouTube Shorts ad under 60 seconds: stitch clips with Timeline
YouTube suggests Shorts ads under 60 seconds. Sume Timeline 1.0 stitches clips into one MP4 at $0.10 per output minute, so 60 seconds costs $0.10.
Written by Sume