OpenAI Agents 0.23 nested agent tools: one Sume key per item
Agents SDK 0.23.0 fixed tool state isolation for nested agents. For paid Sume calls, the idempotency key must still come from your item, not from the agent run.

When one Agents SDK agent calls another as a tool and both can reach Sume, derive the Sume idempotency_key from the item being made, not from the agent run, so a replay or a duplicate call returns the same job. The 0.23.0 fix for nested agent tool state improves isolation in the SDK; it does not know your Sume keys.
The v0.23.0 release (read 2026-10-02) lists, among over 100 fixes, "improved nested agent tool state isolation" and tool identity preservation during async operations. The notes do not describe the exact behaviour, so I treat it as general hardening.
Why does nesting make duplicate paid calls likelier?
A parent agent delegates "make the hero video" to a child agent; the parent times out or re-plans and delegates again. Two child runs now each hold a legitimate reason to call generate_video. Without a shared key they are two paid jobs.
Sume requires idempotency_key on write and paid MCP tools, but its docs describe it as transport and dedup, not human approval (MCP tools and gates). It deduplicates only if both calls send the same key.
What should the key be built from?
From the work item, stable across agents and retries.
| Key built from | Duplicate call result | Verdict |
|---|---|---|
| Order id, shot number, version | Same job returned | Good |
| Agent run id | Different key per run, second paid job | Bad |
| Random UUID per tool call | Second paid job on any replay | Bad |
| Same key, different prompt | 409 idempotency_conflict | Fix the key or the prompt |
How do I make the child agent follow this?
Pass the key in as an argument the parent computes, and tell the child to use it verbatim. Do not let the model invent keys. Then give the child instructions that mirror Sume's own guidance:
- Cap spend with
max_spend_usd, which Sume enforces only when provided. - Wait in slices of at most 55 seconds; a
524onjobs_waitis a transport failure, never a job outcome. - Return the job id to the parent so it never delegates the same shot twice.
Use idempotency_key "order-8823-shot-02-v1" exactly as given.
Call tools_schema for generate_video first. Run dry_run=true
before the paid call. If jobs_wait returns wait_slice_expired,
call jobs_wait again with the same ids. Never create a second job.How do I test duplicate delegation?
Make the parent delegate the same shot twice on purpose, either by calling the child tool twice in a script or by forcing a retry. Both child runs should send the identical key, and Sume should show one job. The second call returns the original job, so the parent can read its status without a new charge.
Then make the two calls differ: same key, changed prompt. You should get 409 idempotency_conflict, and the child should report it instead of looping. A child that answers a conflict by inventing a fresh key has defeated the check.
Finally, run the whole flow under read-only OAuth. The paid tools are hidden, so the child should report that it cannot generate, not improvise. That confirms your instructions do not depend on tools the session cannot see.
Does the SDK retry a paid call for me?
Replay and approval are separate settings in the SDK, covered in Agents SDK approve_unsafe_replay with Sume paid calls. If a replay does reach Sume, a stable key makes it harmless; a changed payload under the same key returns 409 idempotency_conflict. Sume has no way to see which agent made a call, so the dedup lives entirely in the key you pass. Before you ship, read the live contract for every route you call at https://api.sume.com/reference/json, which Sume's docs name as the schema source of truth, and re-read the linked docs pages: limits, scopes and error codes change faster than blog posts do. Treat any number in this post as a snapshot dated 2026-10-02, and prefer the effective fields your own responses return, such as generation_limits, over a static table.
Sources
Related posts
More in Developers
- OpenAI structured outputs: 5000 properties, 10 levels, vs Sume
OpenAI caps strict schemas at 5000 properties and 10 nesting levels. Sume's output_schema uses the same numbers but counts enums and strings differently.
- OpenRouter output_modalities=video vs Sume /v1/videos/models
OpenRouter lists video models three ways. Sume has a /v1/videos/models endpoint with the same shape plus a catalog. Fields to read before you submit.
- OpenRouter video is ZDR-ineligible: what Sume says about retention
OpenRouter says video generation cannot use Zero Data Retention because output is held for retrieval. Sume has no ZDR toggle; what to tell your security team.
- Pace bulk Sume submits with generation_limits, not wave_size_hint
Size in-flight work as concurrency minus active minus queued, capped by queue capacity. A short Node pacer that reads generation_limits from each submit.
Written by Sume