Cloudflare portal Code Mode policy vs Sume script_run

Both shrink tool-call overhead, in different places. What Cloudflare's portal Code Mode policy controls and when Sume script_run is the other half.

5 min readSume
All posts

Cloudflare's Code Mode policy and Sume's script_run solve different halves of the same cost problem. Per Cloudflare, Code Mode policies control how a portal reduces tool definitions and token use (read 2026-10-03), which is about what the model has to read before it calls anything. Sume's script_run is about what happens after: it runs a short JavaScript program on the Sume side that calls tools in a loop, in parallel, or conditionally and returns one value. One shrinks the menu; the other shrinks the number of round trips.

The Cloudflare line is from the portal GA changelog; Cloudflare's page gives no more detail than that sentence, so this post does not describe how the policy works internally. The Sume facts are from MCP tools and gates.

Two layers, two costs

A portal sits between the client and many servers, so it is the place to cut definition overhead across all of them. A single server cannot see the other servers' menus, so it cannot do that. In the other direction, a portal cannot know that your turn needs thirty text-to-speech calls of the same shape; the server's own tools can.

Where each mechanism acts, read 2026-10-03
MechanismActs onSaves
Portal Code Mode policyTool definitions the model loadsTokens spent on menus
Sume script_runThe calls themselvesModel turns and round trips
Sume jobs_wait with job_idsWaiting on many jobsN single waits down to one

What script_run does and does not do

Inside a script, await sume.call(name, arguments) runs any listed tool with the same gates, redaction, and errors as a direct call. Paid creates still need their own idempotency_key. The run is bounded by timeout_seconds (5 to 55), max_calls, and max_paid_calls, and the response returns the script value, a calls[] journal, and the child jobs[] to wait on.

Discovery tools and script_run itself cannot be called from a script, so the script cannot rediscover or recurse. Sume documents it for turns that need three or more independent calls of the same shape.

  • Set max_paid_calls to the number of paid creates you expect, so a loop bug stops early.
  • Give every create inside the loop a distinct idempotency_key, such as one per scene.
  • Hand the returned jobs[] to one batch jobs_wait rather than a wait per id.

An example script body

A script that fans out two image creates with distinct keys. The shape of the arguments depends on the tool, so read tools_schema for generate_image first; this only shows the loop and the key discipline.

const scenes = ["harbor at dawn", "market in rain"];
const out = [];
for (const [i, scene] of scenes.entries()) {
  const r = await sume.call("generate_image", {
    idempotency_key: `scene-batch-a-${i}`,
    max_spend_usd: 1,
    payload: { prompt: scene },
  });
  out.push(r);
}
return out;

Which one first

If the symptom is a huge tool list in context, a portal-level policy is the lever, and you should evaluate it on Cloudflare's documentation. If the symptom is many near-identical paid calls in one turn, script_run is the lever, and a batch jobs_wait with up to 20 job_ids collects the results. When a wait expires with wait_slice_expired, retry the wait with the same ids and never resubmit the paid create. For how clients prompt on parallel calls, see the 2 of 5 prompt post.

Where the two can collide

One risk is worth naming. If a portal policy hides or rewrites tool definitions, a script that calls a tool by name still has to match what the server expects. script_run calls tools by their live names with the same arguments a direct call uses, so keep the contract in view: fetch it with tools_schema from a session that is allowed to see the tool, and keep that call outside the script, since discovery tools cannot run inside one.

A second risk is budget. A script that loops can spend as fast as it can call, which is why Sume gives it max_calls and max_paid_calls as separate bounds. Set both lower than you think you need on the first run, and raise them after you read the calls[] journal. A call that is refused by the gate shows up there with its error, which is a cheaper way to learn than a surprise on the wallet.

Finally, remember the 55-second ceiling on timeout_seconds. A script that creates jobs should return the job ids and let a separate batch wait collect them, rather than trying to wait inside the script.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume