GLM 5.3 Fast or Flash: which one does Sume list?
Sume's agent model list has GLM 5.3 Flash, not GLM 5.3 Fast. What Z.ai's GLM-5.3 page says and what the API's model field accepts.

Sume's model catalog lists GLM 5.3 Flash, routed to Z.ai's glm-5.3-flash through OpenRouter, and it does not list a GLM 5.3 Fast. The row is enabled in the in-app agent model picker behind a feature gate that the catalog calls openrouter. If you search for GLM 5.3 Fast and want it in Sume, the honest answer is that there is no such row; Flash is the nearby entry.
What Z.ai's own page covers
The GLM-5.3 page on Z.ai's docs, read on 2026-10-08, describes a 1M-token context window with a 128K maximum output, always-on reasoning, function calling, streaming, context caching and structured output. It does not mention a Flash or Fast variant, so those names should be checked against your provider's own model list. A third-party host may sell a separate speed-tuned tier under a different name.
| Item | Fact |
|---|---|
| Z.ai GLM-5.3 context window | 1M tokens |
| Z.ai GLM-5.3 maximum output | 128K tokens |
| Z.ai page mentions Flash or Fast | No |
| Sume catalog GLM row | GLM 5.3 Flash (z-ai/glm-5.3-flash via OpenRouter) |
| Sume catalog GLM 5.3 Fast row | None |
| Agent Completions model field | sume-agent only |
What the Sume API accepts
The model picker belongs to Sume's chat. The Agent Completions API takes model as sume-agent only, and a different value returns 400 invalid_request. So you cannot request GLM, Haiku or any other model by name through that endpoint; the Sume agent chooses its own models. If you need GLM 5.3 in your own agent loop, call it from your own client and use Sume's hosted MCP server for the media tools.
Check the catalog behind the gate before promising a user a model. A listed row is enabled for the picker only for workspaces that have the gate.
How to decide
For another retirement-driven catalog change, see Haiku 4.5 retirement and the Sume agent model row.
- Need the model by name in an API call to Sume: not supported;
sume-agentis the only value. - Need GLM in a chat in Sume: look for GLM 5.3 Flash in the picker.
- Need a faster GLM tier: read the vendor or host page for that tier, not this catalog.
- Need cost control on either path:
generation_spend_cap_usdormax_spend_usd.
Why the name matters in a search
People type both names because hosts use different labels for lighter or faster tiers of the same family. A label on one provider's page is not evidence of a model on another's. Sume's catalog entry is a concrete string, z-ai/glm-5.3-flash, which you can check against OpenRouter's listing yourself.
If you are comparing GLM options for an agent, compare on the same terms: context, output limit, tool support and price, each from the provider you will call. Do not carry a number from one host to another.
Sources
Related posts
More in Models
- Does Google keep Omni and Veo prompts for 55 days?
Google's Gemini API page says prompts, context and outputs are kept 55 days for abuse checks. Veo files last 2 days. How that differs from a Sume request.
- Higgsfield Soul on Sume: half a cent per image, batches of 1 or 4
higgsfield-soul costs $0.005 at 720p and $0.0075 at 1080p on Sume. It takes batch sizes 1 or 4, no references. Price table and a four-image call.
- Hy Image 3.5 Preview: 2K or 4K? Tencent and OpenRouter differ
Tencent's launch post says up to 2K; OpenRouter's page lists resolution 1K to 4K. How to test it, and which Sume image rows list a 4K tier or pixel size.
- Hy Image 3.5 Preview: n=1 and seed, against Sume's n and seed
OpenRouter lists Hy Image 3.5 Preview with n capped at 1 and a seed. Sume's Image API takes n up to 4 on most rows and returns 400 for seed. A field table.
Written by Sume