Cloudflare portal Code Mode policy vs Sume script_run
Both shrink tool-call overhead, in different places. What Cloudflare's portal Code Mode policy controls and when Sume script_run is the other half.

Cloudflare's Code Mode policy and Sume's script_run solve different halves of the same cost problem. Per Cloudflare, Code Mode policies control how a portal reduces tool definitions and token use (read 2026-10-03), which is about what the model has to read before it calls anything. Sume's script_run is about what happens after: it runs a short JavaScript program on the Sume side that calls tools in a loop, in parallel, or conditionally and returns one value. One shrinks the menu; the other shrinks the number of round trips.
The Cloudflare line is from the portal GA changelog; Cloudflare's page gives no more detail than that sentence, so this post does not describe how the policy works internally. The Sume facts are from MCP tools and gates.
Two layers, two costs
A portal sits between the client and many servers, so it is the place to cut definition overhead across all of them. A single server cannot see the other servers' menus, so it cannot do that. In the other direction, a portal cannot know that your turn needs thirty text-to-speech calls of the same shape; the server's own tools can.
| Mechanism | Acts on | Saves |
|---|---|---|
| Portal Code Mode policy | Tool definitions the model loads | Tokens spent on menus |
Sume script_run | The calls themselves | Model turns and round trips |
Sume jobs_wait with job_ids | Waiting on many jobs | N single waits down to one |
What script_run does and does not do
Inside a script, await sume.call(name, arguments) runs any listed tool with the same gates, redaction, and errors as a direct call. Paid creates still need their own idempotency_key. The run is bounded by timeout_seconds (5 to 55), max_calls, and max_paid_calls, and the response returns the script value, a calls[] journal, and the child jobs[] to wait on.
Discovery tools and script_run itself cannot be called from a script, so the script cannot rediscover or recurse. Sume documents it for turns that need three or more independent calls of the same shape.
- Set
max_paid_callsto the number of paid creates you expect, so a loop bug stops early. - Give every create inside the loop a distinct
idempotency_key, such as one per scene. - Hand the returned
jobs[]to one batchjobs_waitrather than a wait per id.
An example script body
A script that fans out two image creates with distinct keys. The shape of the arguments depends on the tool, so read tools_schema for generate_image first; this only shows the loop and the key discipline.
const scenes = ["harbor at dawn", "market in rain"];
const out = [];
for (const [i, scene] of scenes.entries()) {
const r = await sume.call("generate_image", {
idempotency_key: `scene-batch-a-${i}`,
max_spend_usd: 1,
payload: { prompt: scene },
});
out.push(r);
}
return out;Which one first
If the symptom is a huge tool list in context, a portal-level policy is the lever, and you should evaluate it on Cloudflare's documentation. If the symptom is many near-identical paid calls in one turn, script_run is the lever, and a batch jobs_wait with up to 20 job_ids collects the results. When a wait expires with wait_slice_expired, retry the wait with the same ids and never resubmit the paid create. For how clients prompt on parallel calls, see the 2 of 5 prompt post.
Where the two can collide
One risk is worth naming. If a portal policy hides or rewrites tool definitions, a script that calls a tool by name still has to match what the server expects. script_run calls tools by their live names with the same arguments a direct call uses, so keep the contract in view: fetch it with tools_schema from a session that is allowed to see the tool, and keep that call outside the script, since discovery tools cannot run inside one.
A second risk is budget. A script that loops can spend as fast as it can call, which is why Sume gives it max_calls and max_paid_calls as separate bounds. Set both lower than you think you need on the first run, and raise them after you read the calls[] journal. A call that is refused by the gate shows up there with its error, which is a cheaper way to learn than a surprise on the wallet.
Finally, remember the 55-second ceiling on timeout_seconds. A script that creates jobs should return the job ids and let a separate batch wait collect them, rather than trying to wait inside the script.
Sources
Related posts
More in Developers
- Cloudflare MCP portal logs: telling Sume read and paid calls apart
Portal logs record tool activity and Logpush exports it. Which Sume tool names mean spend, so a SIEM rule can flag them without reading arguments.
- How to compare AI video models fairly: one prompt, three models, 480p
Submit one prompt to Seedance 2.0 Mini, Wan 3.0 and MiniMax H3 at 480p and 5 seconds on Sume, then judge the clips blind. A runnable Python test harness.
- Join more than 20 voiceover lines: two-level timeline audio concat
Timeline audio concat takes 1 to 20 parts. A 120-line script needs six group joins and one final join, seven jobs, $0.07. A Python planner and offset math.
- Contract-test Sume API responses against openapi.json (pytest)
Validate recorded Sume responses against the OpenAPI schema with jsonschema, including the OpenAPI 3.0 nullable fix. A tested pytest file and fixtures guide.
Written by Sume