GLM-5.3 reasoning cannot be turned off: cap the Sume run instead
Z.ai says GLM-5.3 always reasons, with low, high and max levels. What that means for run time and for a spend cap on a Sume agent run.

Z.ai's GLM-5.3 page says the model always operates with reasoning enabled and supports three reasoning effort levels, low, high and max, and that disabling reasoning is no longer supported. For an agent that calls Sume tools, that means you cannot trade quality for a quick reply by switching thinking off, so cap what matters: wall-clock waits and Sume spend.
What the vendor page says
Reasoning tokens count as output on most providers, but Z.ai's page read today does not give a token price, so check your host for the exact billing.
| Item | Fact |
|---|---|
| Reasoning | Always enabled; cannot be disabled |
| Effort levels | low, high, max |
| Context window | 1M tokens |
| Maximum output | 128K tokens |
| Capabilities listed | Function calling, streaming, context caching, structured output |
| MCP mentioned on the page | No |
Controls that exist on the Sume side
Two controls are independent of the model. The first is the spend cap: max_spend_usd per paid MCP call, which is optional and enforced only when sent, and generation_spend_cap_usd on Agent Completions, which is required. The second is the wait: jobs_wait holds up to 55 seconds per call on the remote server, so a slow reasoning step in your own loop does not extend a Sume call.
A model that always reasons will take longer on the decision turns than on the polling turns. Use the lowest level, low, for status checks, and reserve high or max for choosing tools and building payloads.
A loop that holds up
If you only want Sume's own agent to do the work, the model field accepts sume-agent, and the spend cap is mandatory; see generation_spend_cap_usd is required.
- Set the reasoning level per step,
lowfor polling. - Send
max_spend_usdon every paid call anddry_runbefore the first submit. - Reuse the same
idempotency_keywhen retrying the same intent. - Treat a model timeout as separate from a Sume job state; look the job up before submitting again.
Choosing the level
Z.ai's page names the three levels but this read gives no guidance on which to choose, so treat that as an eval question. Start at low for steps that read state and high for steps that plan, and move to max only where a measured gain justifies the extra time.
The risk to watch is latency. A reasoning model that always thinks will add time to every turn, including polling. Counting turns matters more than raw speed: a loop that polls every 50 seconds spends most of its time waiting on Sume, not on the model.
Sources
Related posts
More in Models
- GLM 5.3 Fast or Flash: which one does Sume list?
Sume's agent model list has GLM 5.3 Flash, not GLM 5.3 Fast. What Z.ai's GLM-5.3 page says and what the API's model field accepts.
- Does Google keep Omni and Veo prompts for 55 days?
Google's Gemini API page says prompts, context and outputs are kept 55 days for abuse checks. Veo files last 2 days. How that differs from a Sume request.
- Higgsfield Soul on Sume: half a cent per image, batches of 1 or 4
higgsfield-soul costs $0.005 at 720p and $0.0075 at 1080p on Sume. It takes batch sizes 1 or 4, no references. Price table and a four-image call.
- Hy Image 3.5 Preview: 2K or 4K? Tencent and OpenRouter differ
Tencent's launch post says up to 2K; OpenRouter's page lists resolution 1K to 4K. How to test it, and which Sume image rows list a 4K tier or pixel size.
Written by Sume