FLUX 3 Image bounding boxes vs Sume masked edit: what each takes
FLUX 3 Image takes a JSON list of boxes at the end of the prompt. Sume's image edit takes a mask image: mask_image_url with image_urls. Differences, no overlap.

On Sume you change one region of an image with a mask image, not boxes in the prompt. Send image_urls plus mask_image_url on the Image 1.0 route. FLUX 3 Image works differently: its release notes say to add a JSON list of boxes to the end of prompt, with no separate parameter. Sume's docs list no FLUX 3 model id.
BFL's side is from its release notes and Sume's from the Image 1.0 and Image generation docs, read 2026-10-01.
How do bounding boxes work in FLUX 3 Image?
BFL describes placing every element of a layout, such as a headline, a face in a crowd or each panel of a grid, by boxes appended to the prompt. It also lists pixel-exact local edits: mark the element to change and keep the rest. Its docs link a separate bounding-box tutorial for the exact box format, which this post does not reproduce.
What does Sume take to change one region?
A mask. The Image 1.0 table lists a masked edit as image_urls plus mask_image_url, described as a mask image URL for edit flows. The ChatGPT Image 2.5 models on the Image Router use mask_url instead, so the field name depends on the route.
| Side | Region marked by | Where |
|---|---|---|
| FLUX 3 Image | JSON list of boxes | End of prompt |
| Sume Image 1.0 route | Mask image | mask_image_url with image_urls |
| Sume ChatGPT Image 2.5 | Mask image | mask_url |
How do I send a masked edit?
Host the source and the mask at public HTTPS URLs and send both. The Image 1.0 docs say new integrations should use POST /v1/images and that Image 1.0 is retiring soon, so check the mask field for your route. The prompt says what to put in the marked area.
curl -X POST https://api.sume.com/v1/image-1.0/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: image-mask-001" \
-d '{
"prompt": "Replace the masked sign with a plain wooden board",
"image_urls": ["https://example.com/street.png"],
"mask_image_url": "https://example.com/street-mask.png"
}'Can I convert boxes to a mask?
Sume's docs do not describe a boxes input, so you would draw the mask yourself, one filled rectangle per box, before upload. The docs do not state a mask colour convention here; check the mask field notes and test on one image first.
Sources
Related posts
More in Models
- FLUX 3 x mimic: one backbone for video, audio and robot actions
BFL says FLUX 3 is one model trained jointly on images, video and audio, and FLUX-mimic adds robot actions. What that does and does not tell a media API user.
- FLUX 3 video continuation caps at 15 s: extending a clip on Sume
BFL cut FLUX 3 video continuation (v2v) to 15 seconds on Aug 17; other modes stay at 20. Sume has no continuation mode: chain a last frame instead.
- FLUX 3 Video 4K uhd (3840x2176) and which Sume video models do 4K
FLUX 3 Video's uhd resolution is 3840 x 2176 at 16:9. On Sume, a 4K request goes to a model whose supported_resolutions lists it; here is what the docs say.
- Flux TTS expressivity -2 to 2 vs Sume's emotion guide
Deepgram's Flux TTS expressivity runs -2 to 2 (0 nominal). Sume TTS has no such dial: generation_config takes volume, speed and a free-text emotion guide.
Written by Sume