Three image edits at once or one by one? Cost and drift on Sume
FLUX 3 Image edits several boxes in one pass. On Sume, one prompt with three numbered changes costs one edit; three chained edits cost three and can drift.

Black Forest Labs' FLUX 3 Image docs describe element kinds (new, move, anchor) with source and target boxes, so several elements change in one edit, and they warn that pixels outside the boxes only usually stay the same: shadows, reflections or nearby lighting may still change (Black Forest Labs docs, read 2026-10-04). The Sume Image API has no box field. The buyer question is the same, though: should three changes go in one request or three?
The cost side is simple
Each request is billed per image, so chaining multiplies cost. At Nano Banana 2's 1K rate of about $0.10 per image (catalog list $0.08 × 1.25), one call that asks for three changes is about $0.10, and three chained calls are about $0.30. A failed generation is not billed, so a retry costs only when it succeeds (Sume Image API).
| Approach | Requests | Approx. cost | Risk |
|---|---|---|---|
| One prompt, three numbered changes | 1 | $0.10 | Model may skip or blend a change |
| Three chained edits, one change each | 3 | $0.30 | Each pass can drift the untouched area |
| One call, then a fix-up for the missed change | 2 | $0.20 | Only the missed change is re-run |
The drift side needs a test
Neither approach is better by rule. A single pass leaves fewer chances to shift lighting; a chain gives each change its own full attention. Test it on ten of your own photos, count how many results kept all three changes, and diff the untouched areas. If you chain, save PNG between passes so JPEG artifacts do not stack.
A single-call prompt
Number each change and end with a keep-the-rest sentence. If a change goes missing, re-run only that change on the result.
Make exactly three changes to the reference photo:
1. Change the left cushion to mustard yellow.
2. Remove the cable on the floor.
3. Replace the picture on the wall with a plain white frame.
Keep everything else identical: layout, lighting, shadows, perspective.
Do not add any text or markings.Sources
Related posts
More in Models
- Tidal stopped paying wholly AI tracks: what a video score is
Tidal stopped paying royalties on wholly AI tracks on July 15, 2026. What that means for a video soundtrack, and how Sume's Music Router fits.
- Trending TikTok metadata to a Seedance 2.5 clip
Sume's trending-videos search returns TikTok metadata for $0.10 a call, not video files. Use it to write a brief, then render an original Seedance 2.5 clip.
- Veo 3.1 extend adds 7 s up to 20 times, 720p only: plan a long clip
Google's Veo docs let you extend a clip by 7 seconds up to 20 times, at 720p only. Read the arithmetic, then compare it with one-call 30 second models on Sume.
- Veo 3.1 Lite 1,024-token prompt limit vs Omni's 20,000 characters
Google's Veo 3.1 Lite page caps text input at 1,024 tokens. Sume's gemini-omni-flash-1.1 row takes 20,000 characters. What that means for long shot lists.
Written by Sume