Wall art in a room: preview video from one photo and a first frame
Turn one room photo and an artwork file into a short walk-up video: one Sume image edit for the first frame, then image-to-video. Both calls and the cost.

A still of a print above the sofa answers one question. A five-second push toward the wall answers the next: how does it feel in the room? You can get both from one photo of the room and the artwork file, with the still acting as the video's first frame.
Step 1: hang the art in the room photo
Send the room photo first and the artwork second in input_references. Sume lists ChatGPT Image 2.5 as openai/gpt-image-2.5 (Flare) and openai/gpt-image-2.5-sunburst, with up to 16 references. OpenAI's guide says to choose Sunburst where editing precision matters most, so this edit uses Sunburst. Use aspect_ratio: "auto" so the frame matches the photo; omitting it is not the same as auto.
In Sume's golden-amount fixture, a high-quality Sunburst edit with one reference bills $0.29. Treat that as an estimate, since input tokens change it.
Step 2: use that image as the first frame
On POST /v1/videos, frame_images with frame_type: "first_frame" makes the request image-to-video. minimax-h3-max takes a first frame, runs 5 to 15 seconds, and bills $0.10 per second at 768p on Sume. A five-second clip is $0.50. The model always produces native stereo audio and has no toggle, so plan to mute it in your player or timeline.
import os
import requests
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
edit = requests.post(
"https://api.sume.com/v1/images",
headers=H,
json={
"model": "openai/gpt-image-2.5-sunburst",
"prompt": "Image 1 is a room, image 2 is framed art. Hang image 2 "
"centered above the sofa. Change nothing else.",
"aspect_ratio": "auto",
"input_references": [
{"type": "image_url", "image_url": {"url": u}}
for u in ("https://example.com/room.jpg",
"https://example.com/art.jpg")
],
},
timeout=60,
)
if edit.status_code != 200:
raise SystemExit(f"edit returned {edit.status_code}: poll the job")
frame = edit.json()["data"][0]["url"]
print(frame)Step 3: ask for a slow move and nothing else
Submit the video with the frame and a prompt that moves the camera, not the room: a slow push-in toward the wall, daylight steady, no people, art and furniture unchanged. Then poll GET /v1/videos/{id} until completed, as the video docs describe.
| Step | Model | Basis | Cost |
|---|---|---|---|
| First-frame image | openai/gpt-image-2.5-sunburst | 1 edit, high quality, fixture estimate | About $0.29 |
| Walk-up clip | minimax-h3-max, 768p | 5 s at $0.10 per second | $0.50 |
| Total | About $0.79 |
Check the frame before you pay for motion
The edit costs under a third of the clip. Compare the art in the still with your file at full size first; if the print was redrawn or the room changed, rerun the edit, not the video. A failed image call is not billed, and the docs say a completed one is billed in full.
Remember the render is a concept, not a scale drawing. Measure the real wall before you order.
Sources
Related posts
More in Use cases
- Walmart Recognized Reviewer: under 15 reviews, 70% content score
Walmart's Recognized Reviewer now covers items under 15 reviews, if the content quality score is 70% or higher. What to fix first, and what Sume can make.
- Was-price in a sale video: FTC former-price rule and cues
A crossed-out 'was' price on screen needs a real former price. What FTC 16 CFR 233.1 says and how to burn it as one exact caption cue on Sume for $0.20.
- Webinar reminders: three 20-second avatar clips per event
Three 20-second 16:9 avatar reminders per webinar cost $11.04 on Sume's standard tier, $14.70 on plus. Twelve webinars a year is $132.48 on standard.
- Wedding photographer's engagement sneak-peek reel: $0.10 each on Sume
A 24-second sneak-peek reel from 8 delivered photos costs $0.10 to render on Sume, plus one $0.125 music bed reused across couples: $1.125 for ten.
Written by Sume