Idempotency-Key for Agent Completions: reuse the model's tool call id
Retries after a timeout must not start a second paid Sume run. Derive Idempotency-Key from the tool call id your model returned, such as the OpenAI call_id.

Use one Idempotency-Key per intended run, and derive it from an identifier the model already gave you, such as the function call's call_id in the OpenAI Responses API. If your server retries after a timeout, Sume returns the original receipt with idempotency_hit: true instead of starting and billing a second run. Reuse the same key with a different payload and you get 409 idempotency_conflict.
What each side guarantees
The Agent Completions page says Idempotency-Key behaves as it does on Action runs: a replay returns the original receipt flagged idempotency_hit: true, and a different payload under the same key returns 409 idempotency_conflict. On the OpenAI side, the function calling guide describes the model returning function_call items that carry a call_id, which you answer with a function_call_output item. That id is stable for the call, which makes it a natural key.
Choosing the key
| Key source | Retry-safe? | Risk |
|---|---|---|
Model tool call_id | Yes, same call same id | None if payload is unchanged |
| Random UUID per attempt | No | Every retry is a new paid run |
| Hash of the instruction text | Mostly | Two intended runs with identical text collapse into one |
| Timestamp | No | Differs on every retry |
Handling the two replies
Treat idempotency_hit: true as success and carry on with the run id you were given; do not start a fresh run. Treat 409 idempotency_conflict as a bug in your code, not a transient error: the same call id is paired with a changed payload, usually because your handler rewrote the cap or the instruction between attempts. Fix the handler so the payload is built once per call id, then retried unchanged.
Return the run id to the model as the tool output, with a note that the work is asynchronous, so the model polls rather than starting another.
- Build the request body once, store it with the key, and retry that exact body.
- Clamp the cap before computing the body, not after.
- Log key and run id together.
- Keep one key per intended run, never one per attempt.
Beyond create
The documented read and cancel calls are sent without a key; only the create call starts a run. If you also fire runs from a schedule or a Format, check how those routes treat keys on their own pages; the same name does not guarantee the same window. See Safe automation for the general guardrails.
Sources
Related posts
More in Developers
- Agent Completion messages[]: system and user turns become one prompt
How Sume joins messages[] into one prompt, what a system turn can and cannot do, and why a GPT-6.1 Sol or Sonnet 5.5 chat history cannot be replayed as is.
- Agent Completion output_schema: fail a CI build when it is invalid
Bind output_schema to a Sume Agent Completion and gate CI on the receipt: status completed, output present, output_error empty. Strict schema rules explained.
- Agent Completions 403 insufficient_scope: an older key needs replacing
POST /v1/agent/completions returns 403 insufficient_scope for a key that predates Agent Completions or a service-account key. How to tell which and replace it.
- Lazy-load placeholder for an AI image: average color from Sume
Compute the average color of a generated image with Pillow, use it as the background of the image box and avoid a white flash while the real file loads.
Written by Sume