Limit an agent's paid MCP calls with script_run max_paid_calls
script_run runs a short program on the Sume side with max_calls, max_paid_calls and a 5-55 second timeout. What it can bound, and what it does not cap.

script_run on Sume's hosted MCP runs a short JavaScript program that calls other Sume tools in a loop, in parallel, or with conditions, and returns one value. Its limits are timeout_seconds (5 to 55), max_calls, and max_paid_calls. It is a way to bound how many paid calls one agent turn can make, but it is not a dollar cap.
What the docs say
The tools and gates page recommends script_run when a turn needs three or more independent calls of the same shape, such as one tts_create per sentence or one generate_image per scene. Inside the script, await sume.call(name, arguments) runs any listed tool with the same gates, redaction, and errors as a direct call.
| Property | Value |
|---|---|
| timeout_seconds | 5 to 55 |
| max_calls | Set per run |
| max_paid_calls | Set per run |
| Paid creates inside | Each still needs its own idempotency_key |
| Return value | The returned value, a calls[] journal, and child jobs[] for jobs_wait |
| Not callable from a script | Discovery tools, and script_run itself |
Why it helps with a cheap, fast agent model
Fast models such as Claude Haiku 5.5, which Anthropic's page (read 2026-10-08) pitches for subagents and high-volume work, will happily take many small steps. Each step pays prompt tokens again. Moving a fan-out of ten identical tool calls into one script means the model pays for one call and one returned value instead of ten results in its context.
The count limits then bound the blast radius of a bad plan. If the model decides to loop over 400 scenes by mistake, max_paid_calls stops the script at the number you set.
What it does not do
There is no dollar field in the list above. For a dollar ceiling, use the gates the docs name elsewhere: max_spend_usd on a paid tool, which Sume enforces only when you provide it, and dry_run=true for an estimate that does not submit. Wallet balance and generation admission still apply on every paid call.
Also note the scope rule. Under OAuth mcp:read the write and paid tools are hidden, and the docs say there is no separate paid scope; you need mcp:write or an API key.
A practical pattern
The points that matter here, in the order you will hit them:
- Call
generation_admission_previewor adry_runfirst for any burst. - Set
max_paid_callsto the number of items you intend, not a round large number. - Pass a stable
idempotency_keyper item so a retried script does not double-submit. - Wait on the returned
jobs[]withjobs_waitrather than re-running the script.
Choosing the numbers
Suppose a script renders 12 scenes, then waits for the jobs. That is 12 paid creates. Setting max_paid_calls to 12 means a thirteenth paid call stops the script. max_calls has to cover the paid creates plus any reads, so a value of 14 would allow two reads. The timeout is separate: 5 to 55 seconds, so long generation work is not waited on inside the script. The docs return child jobs[] so you can use jobs_wait afterward.
Sources
Related posts
More in Developers
- result_ready vs terminal vs completed: gate the Sume result fetch
Poll a Sume job until terminal, fetch the result only when result_ready is true. Failed and canceled jobs answer 409 job_not_completed on /result.
- Sume run webhook retries: ten attempts span about 3 hours
Ten delivery attempts with 30-second doubling backoff capped at one hour add up to 11,010 seconds before jitter. The timeline, and what to do after exhaustion.
- Sume run webhook says OK but output is null: handle degraded
A Sume run can complete, bill you, and still have output null. Branch on outcome, not status, and read artifacts and output_error before you retry.
- Sume submit: request_id is the job id you poll (curl, jq)
After an async submit, data.request_id is the id for GET /v1/jobs/:id. error.request_id is for support tickets. A curl and jq check.
Written by Sume