FLUX 3's element table as app state: run it on Sume with job metadata
Keep a FLUX 3 style element table in your own app, build each Sume edit from it, and tag every job with the table version through the metadata field.

You can get most of FLUX 3 Image's "element table as project state" workflow on Sume by keeping the table in your own app: a dictionary of named elements with a description, a box and a locked flag, from which you build each edit request. Sume does not store the table, and no Sume model takes boxes, but a job's metadata field lets you stamp every generation with the element and table version that produced it.
BFL's FLUX 3 Image documentation describes elements as New, Anchor or Move, each with a box on a 0 to 1000 grid, and the idea of an agent that carries the list from turn to turn. The pattern below is the same idea with the parts Sume provides: a prompt, references, a mask and metadata.
What the app owns and what Sume owns
Be explicit about who holds what, because a stateless API forgets everything between calls. Sume's Image API takes a prompt, input_references and an optional mask_url, and returns a hosted URL. Everything that makes the edit repeatable lives on your side.
| Concern | Lives in | Field or tool |
|---|---|---|
| Element descriptions and boxes | Your app | Element table |
| Which elements are locked | Your app | locked flag, turned into a keep list |
| Mask for the element being edited | Your app, hosted at a public HTTPS URL | mask_url (ChatGPT Image 2.5 only) |
| Source image for this step | Your app | input_references |
| Which step produced a job | Both | metadata, stored on the job, not sent to the provider |
The table class
request_for builds the Sume payload for one element: it says to change the masked area only, lists every locked element as something to keep, and attaches metadata with the element key, a version counter and a short digest of the table. That digest is the useful part later: two jobs with the same digest came from the same table state. The metadata is stored on the job and not forwarded to the model, per the Image API docs, so it is safe to put internal labels there.
Keep values small and non-sensitive anyway. The field exists for you to find a job again, not to hold content.
import hashlib
import json
class ElementTable:
"""Your app's source of truth for an image: boxes on a 0-1000 grid, [y0, x0, y1, x1]."""
def __init__(self):
self.items = {}
self.version = 0
def set(self, key, desc, box, locked=False):
self.items[key] = {"desc": desc, "box": box, "locked": locked}
self.version += 1
def digest(self):
blob = json.dumps(self.items, sort_keys=True).encode()
return hashlib.sha256(blob).hexdigest()[:12]
def request_for(self, key, input_url, mask_url):
el = self.items[key]
keep = [v["desc"] for k, v in self.items.items() if k != key and v["locked"]]
prompt = f"In the masked area only: {el['desc']}."
if keep:
prompt += " Keep unchanged: " + "; ".join(keep) + "."
return {
"model": "openai/gpt-image-2.5",
"prompt": prompt,
"input_references": [{"type": "image_url", "image_url": {"url": input_url}}],
"mask_url": mask_url,
"aspect_ratio": "auto",
"metadata": {"element": key, "table_version": self.version, "table_digest": self.digest()},
}
table = ElementTable()
table.set("logo", "the white logo on the cap", [120, 380, 220, 620], locked=True)
table.set("mug", "a red enamel mug", [560, 600, 900, 820])
print(json.dumps(table.request_for("mug", "https://example.com/in.png", "https://example.com/mask.png"), indent=1))Running a step
To run it, host the source image and the mask, post the payload to /v1/images as in the other posts in this series, and read data[0].url. A 200 carries the image; a 202 carries a job envelope, so branch on the status code, as the docs say. When you get the result, save it as the new source and bump the table version before the next element.
Because the table is yours, undo is cheap: restore an earlier table version and the source image that went with it. The same habit of saving every result also helps with the multi-round edit workflow.
Limits to keep in mind
A mask is guidance on Sume, not a clip. Run a check for changes outside the box after each step, as in checking edits with numpy, and keep the keep-list wording in every request, because repeated edits drift when it is dropped. Only ChatGPT Image 2.5 accepts mask_url; other models return 400 unsupported_parameter if you send it.
The fields used here are in the Image API docs; metadata behavior on the retiring alias is described on the Image 1.0 page.
Sources
Related posts
More in Developers
- FLUX 3 Image's New, Anchor and Move elements as Sume edit prompts
BFL's FLUX 3 Image tags every box as New, Anchor or Move. Sume has no box field, so here is how each tag maps to a prompt, a mask or a two-pass edit.
- FLUX 3 Image on OpenRouter: n=1, seed, base64 vs Sume
OpenRouter lists FLUX.3 Image with one image per call, a seed and base64 PNG output. How each differs from Sume's POST /v1/images, where FLUX 3 is not listed.
- FLUX 3 Image on Replicate: safety_tolerance 0-4 vs Sume
Replicate's FLUX 3 Image form has safety_tolerance 0-4, grounding, output_quality and 768sq-4k. Which of those inputs Sume's image API has, and what it returns.
- Format contents PUT 409 format_content_sha_required: send the sha
PUT on a Format file that already exists needs its blob sha. Read the sha with GET, retry, and tell the three 409 codes apart from the package If-Match guard.
Written by Sume