Format run input is data, not instructions: a boundary, not a sandbox

Format runs write input to a file as data, not as instructions. That helps against prompt injection but is not a sandbox. The spend cap bounds the damage.

3 min readSume
All posts

What the docs promise

When you call a Sume Format with an input object, the whole object is written to /workspace/inputs/sume-action-input.json and treated as data. Your instruction is what the agent is told to do. Customer text placed in input does not become part of the instruction.

That is a trust boundary, not a sandbox. A model can still read the file and be influenced by what it says, so the separation lowers risk without removing it.

Where to put what

The table shows the habit this suggests.

Where each kind of text belongs in a Sume run, read 2026-10-08
Kind of textFieldReason
What the agent should doinstructionOnly the first ~4000 characters reach the agent
Order data, product copy, user textinputWritten as data to a file; 64 top-level keys, 2 MiB
Images from customersattachmentsUp to 30 images, 30 MB each
The most you will spendgeneration_spend_cap_usdRun-level ceiling, max $500

Why a cap matters more than a filter

No filter catches every injection. Assume one eventually gets through and ask what it could do. For a Format run, the worst case is bounded by generation_spend_cap_usd and by the tools the Format has. A $20 cap on a product-copy run means a hijacked run cannot spend more than $20.

Agent Completions needs this cap on every call, because it has no default. Set it from the job, not from a global constant.

Controls to add

Keep the instruction fixed and versioned in your code. Pass only the fields the Format needs in input. Strip secrets before you build the object, since the whole thing is written to a file. Use on_active_run: "reject" where a second run on the same object would be a problem.

Log request ids, job ids and statuses. Do not log API keys, signed URLs or raw private URLs.

  • One workspace per key: tools should not accept a workspace id from the model.
  • Use a webhook to receive results, not the model's reply as a command.
  • Review the outputs before anything is published.

A concrete example

Say a support team sends a customer's message to a Format that drafts a reply image. The message goes in input.customer_message. The instruction says what to draw and mentions the field by name. If the message contains ignore your instructions, the agent still reads it as a field, but nothing stops a model from being swayed.

So pair the boundary with a small cap, a Format limited to the tools it needs, and review before sending anything to a customer.

  • Cap per run, set by the use case.
  • Review step before publish.
  • Different Formats for different trust levels.

Keep a written record

Document the decision in the repository next to the code that makes the call, so the next engineer sees why the choice was made and which docs page it came from. Re-read that page when you upgrade a client or change a key, since gates and limits are the parts most likely to differ from what you remember.

A short note of the date you last verified the behaviour, such as 2026-10-08, is enough for a reviewer to know how fresh the claim is.

Sources

Related posts

More in Agents

All Agents posts

Written by Sume