GLM-5.3 reasoning cannot be turned off: cap the Sume run instead

Z.ai says GLM-5.3 always reasons, with low, high and max levels. What that means for run time and for a spend cap on a Sume agent run.

5 min readSume
All posts

Z.ai's GLM-5.3 page says the model always operates with reasoning enabled and supports three reasoning effort levels, low, high and max, and that disabling reasoning is no longer supported. For an agent that calls Sume tools, that means you cannot trade quality for a quick reply by switching thinking off, so cap what matters: wall-clock waits and Sume spend.

What the vendor page says

Reasoning tokens count as output on most providers, but Z.ai's page read today does not give a token price, so check your host for the exact billing.

GLM-5.3 reasoning facts, Z.ai docs, read 2026-10-08
ItemFact
ReasoningAlways enabled; cannot be disabled
Effort levelslow, high, max
Context window1M tokens
Maximum output128K tokens
Capabilities listedFunction calling, streaming, context caching, structured output
MCP mentioned on the pageNo

Controls that exist on the Sume side

Two controls are independent of the model. The first is the spend cap: max_spend_usd per paid MCP call, which is optional and enforced only when sent, and generation_spend_cap_usd on Agent Completions, which is required. The second is the wait: jobs_wait holds up to 55 seconds per call on the remote server, so a slow reasoning step in your own loop does not extend a Sume call.

A model that always reasons will take longer on the decision turns than on the polling turns. Use the lowest level, low, for status checks, and reserve high or max for choosing tools and building payloads.

A loop that holds up

If you only want Sume's own agent to do the work, the model field accepts sume-agent, and the spend cap is mandatory; see generation_spend_cap_usd is required.

  • Set the reasoning level per step, low for polling.
  • Send max_spend_usd on every paid call and dry_run before the first submit.
  • Reuse the same idempotency_key when retrying the same intent.
  • Treat a model timeout as separate from a Sume job state; look the job up before submitting again.

Choosing the level

Z.ai's page names the three levels but this read gives no guidance on which to choose, so treat that as an eval question. Start at low for steps that read state and high for steps that plan, and move to max only where a measured gain justifies the extra time.

The risk to watch is latency. A reasoning model that always thinks will add time to every turn, including polling. Counting turns matters more than raw speed: a loop that polls every 50 seconds spends most of its time waiting on Sume, not on the model.

Sources

Related posts

More in Models

All Models posts

Written by Sume