Farm stand weekend reel: GPT Image 2.5 and Wan 3.0 clips for $1.41
Four 9:16 first frames from GPT Image 2.5 and four 5-second Wan 3.0 clips at 480p cost $1.41 with a Timeline render on Sume. The steps and the checks.

A farm stand can make a four-shot weekend reel for $1.41 on Sume: four 9:16 first frames from GPT Image 2.5 at medium ($0.0555 in all), four 5-second Wan 3.0 clips at 480p ($1.25) and one Timeline render of under a minute ($0.10). At 720p the same reel is $2.66. The images fix what the pumpkins and the hay look like, so the clips animate a picture you already approved instead of inventing a scene.
The price build-up
Prices are Sume's quotes, which add the 1.25 house margin to the provider list. The image size is 1152x2048, a 9:16 frame whose edges are multiples of 16, as the Image API page requires. Wan 3.0 accepts 2 to 30 seconds per the video docs, and its price scales with seconds and resolution.
| Step | Unit price | Count | Subtotal |
|---|---|---|---|
| GPT Image 2.5, medium, 1152x2048 | $0.0139 | 4 | $0.0555 |
| Wan 3.0, 480p, 5 s | $0.3125 | 4 | $1.2500 |
| Timeline render, up to 1 min | $0.10 | 1 | $0.1000 |
| Total at 480p | $1.41 | ||
| Total at 720p (clips $0.625 each) | $2.66 |
Step 1: four frames
Write four prompts that share a look: the same hay-bale table, the same morning light, one subject each, such as pumpkins, apple crates, cider jugs and the hand-painted sign. Render at low first ($0.0060 each) and keep the best of each, then re-render the winners at medium. Nothing here depends on old image ids: OpenAI lists gpt-image-1 for shutdown on 2026-10-23, and openai/gpt-image-2.5 is the current id on Sume (OpenAI, read 2026-10-05).
Order the shots as a story: wide view of the stand, a close product shot, a person-free detail of the sign and prices, then the address card. Four shots at 5 seconds is a 20-second reel, and you could add more shots later at about 31 cents each at 480p.
Step 2: check Wan before you submit
The video catalog lists each model's supported_aspect_ratios, supported_durations and supported_frame_images. Read them for wan-3.0 instead of guessing; a ratio or frame type the model does not list is not safe to send. This script prints that row.
import os
import requests
r = requests.get(
"https://api.sume.com/v1/videos/models",
headers={"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"},
timeout=30,
)
r.raise_for_status()
for m in r.json()["data"]:
if m["id"] == "wan-3.0":
for key in ("supported_resolutions", "supported_aspect_ratios",
"supported_durations", "supported_frame_images"):
print(key, m.get(key))Step 3: animate and assemble
Each clip is a POST /v1/videos with model: "wan-3.0", duration: 5, resolution: "480p" and your image in frame_images with frame_type: "first_frame". The prompt should describe only motion, such as steam rising from the cider or the camera drifting along the crates. Use one Idempotency-Key per shot so a retry never buys a second clip. Then place the four clips on a Timeline 1.0 render, which bills $0.10 per output minute rounded up.
Add a voice line or a music bed on the Timeline audio spine rather than relying on clip audio, so the hours and the address are spoken once, clearly. Put the opening hours in a caption you control; do not trust a video model to draw them.
Name the files by shot and date, such as stand-w41-shot2.mp4, and keep the job ids next to them. If shot three fails or looks wrong, you rerun only that 5-second clip at 31 cents, not the reel, and the Timeline render is the last step you run.
Why a first frame and not text-to-video
A text-only prompt asks the video model to invent the stand, and four separate prompts give four different stands. A first frame makes the picture the starting point, so the pumpkins, the sign and the light match from shot to shot. The frame also costs less than a retry: a medium frame is under 1.4 cents, while a rejected 5-second clip at 480p is 31 cents.
That ratio sets the order of work. Spend your iterations on the cheap still, approve it, and only then pay for motion. If you plan a voice line, write it first and pick shot lengths to fit it, rather than cutting the voice to the clips afterward.
Where it can go wrong
A 480p clip is soft on a large phone, so look at one before you buy all four, and move to 720p if the sign text is blurry. A first frame with legible lettering can drift as the clip plays. Keep signs out of the motion shots, or crop to the sign in the last second. Provider prices can change, so read usage.cost on the first job and update this table.
Sources
Related posts
More in Use cases
- Fashion lookbook video in 3:4 or 9:16: which Sume models take it
Wan 3.0, Seedance and MiniMax H3 take 3:4 and 9:16 on Sume; Kling and Omni do not take 3:4. Wan at 720p is $0.125 a second, H3 at 768p is $0.075.
- Feedback request video for an NPS survey: AI avatar script, cost
A 10-second avatar clip that asks for survey feedback: what to say, what to leave to the email, and the cost for 1,000 sends if you render per segment.
- Festive caption colors for holiday ads: slam style design overrides
TikTok's holiday guide asks for bold captions and festive overlays at peak. One caption job with design colors burns a red-and-green slam line for $0.20.
- 15 listing tours in five languages: 75 caption jobs, $15.00
Burning translated room labels on 15 listing tour videos in Spanish, Portuguese, French, German and Italian is 75 Sume caption jobs at $0.20 each, $15.00.
Written by Sume