GLM-5.3 off-peak 50% points vs a Sume dollar spend cap
Z.ai meters GLM-5.3 in points with a 50% off-peak rate. Why a Sume dollar cap is a different meter, and how to schedule agent work.

Z.ai's GLM-5.3 page describes a points-based quota system in which model calls made during off-peak hours, including all day on weekends, consume 50% of the standard points. A Sume spend cap is in US dollars on Sume's generation wallet, so the two never offset each other. Scheduling your agent's reasoning for off-peak hours saves Z.ai points. It does not reduce what a Sume render costs.
Two meters, two caps
The Z.ai page does not give a dollar conversion for points, so the table does not either.
| Meter | Unit | Control |
|---|---|---|
| Z.ai GLM-5.3 quota | Points; off-peak and weekends use 50% | Schedule calls off-peak |
| Sume paid MCP call | USD | max_spend_usd, optional, enforced when sent |
| Sume Agent Completion | USD | generation_spend_cap_usd, required |
| Sume write or paid call retry | None | Reuse idempotency_key |
Using the discount without a surprise
A batch of Sume jobs can wait for off-peak: queue the intent, then run the planning model when the cheaper rate applies. Sume's schedules and Agent Completions help here. An Agent Completion runs the Sume agent on an ad-hoc prompt, and a schedule stores what to do on a timer. In both, Sume's own agent chooses its models; you cannot pick GLM by name through the API.
For an agent loop you run yourself with GLM-5.3, the loop can run at night and submit to Sume's hosted MCP server with a cap on each call. The render itself is billed by Sume at the time it runs, whatever the GLM time of day.
A guard to keep
The same split between a vendor's peak pricing and Sume's meter is in DeepSeek peak hours and Sume spend caps.
- Do not let a cheap-hours loop run unattended without a per-call
max_spend_usd. - A retry loop that waits for off-peak must keep one
idempotency_keyper intent so the same job is not created twice. - Separate dashboards: Z.ai points on one side, Sume spend on the other.
Reading the fine print on your plan
Quota rules differ by plan and can change. The GLM-5.3 page states the off-peak rate, but it does not define the peak window in the text read here, so check your plan page for the hours before you schedule anything around them.
Treat the discount as a nice-to-have, not as a design constraint. The render, the part with a dollar cost, runs when you submit it, and Sume's jobs_wait and receipts do not depend on the hour.
Sources
Related posts
More in Pricing
- GPT Image 2.5 on fal runs $0.00402 to $0.40026 per image: on Sume
fal's ChatGPT Image 2.5 Flare page lists $0.00402 for 1024x768 low and $0.40026 for 3840x2160 max. At Sume's 1.25 multiplier that is about $0.005 to $0.50.
- Does gpt-image-2.5 cost more at 4K? Prices from 1024 to 3840x2160
gpt-image-2.5 price by pixel size on Sume: 1024x1024, 1536x1024, 2048x2048 and 3840x2160 at medium and high, plus the custom-size rules from the docs.
- gpt-image-2.5 xhigh and max at 2048 and 4K: up to $0.54 an image
gpt-image-2.5 max costs $0.5353 at 2048x2048 and $0.5004 at 3840x2160 on Sume; xhigh is $0.2379 and $0.2224. Size-by-tier table and a spend guard in Python.
- Grok Imagine video 1.5: 1.25 cents a second at any resolution
grok-imagine-video-1.5 bills $0.0125 a second on Sume at 480p or 720p. A table of 4 to 15 second clips and what 100 and 1,000 of them cost.
Written by Sume