GLM-5.3 off-peak 50% points vs a Sume dollar spend cap

Z.ai meters GLM-5.3 in points with a 50% off-peak rate. Why a Sume dollar cap is a different meter, and how to schedule agent work.

5 min readSume
All posts

Z.ai's GLM-5.3 page describes a points-based quota system in which model calls made during off-peak hours, including all day on weekends, consume 50% of the standard points. A Sume spend cap is in US dollars on Sume's generation wallet, so the two never offset each other. Scheduling your agent's reasoning for off-peak hours saves Z.ai points. It does not reduce what a Sume render costs.

Two meters, two caps

The Z.ai page does not give a dollar conversion for points, so the table does not either.

Points vs dollars, Z.ai docs and Sume docs, read 2026-10-08
MeterUnitControl
Z.ai GLM-5.3 quotaPoints; off-peak and weekends use 50%Schedule calls off-peak
Sume paid MCP callUSDmax_spend_usd, optional, enforced when sent
Sume Agent CompletionUSDgeneration_spend_cap_usd, required
Sume write or paid call retryNoneReuse idempotency_key

Using the discount without a surprise

A batch of Sume jobs can wait for off-peak: queue the intent, then run the planning model when the cheaper rate applies. Sume's schedules and Agent Completions help here. An Agent Completion runs the Sume agent on an ad-hoc prompt, and a schedule stores what to do on a timer. In both, Sume's own agent chooses its models; you cannot pick GLM by name through the API.

For an agent loop you run yourself with GLM-5.3, the loop can run at night and submit to Sume's hosted MCP server with a cap on each call. The render itself is billed by Sume at the time it runs, whatever the GLM time of day.

A guard to keep

The same split between a vendor's peak pricing and Sume's meter is in DeepSeek peak hours and Sume spend caps.

  • Do not let a cheap-hours loop run unattended without a per-call max_spend_usd.
  • A retry loop that waits for off-peak must keep one idempotency_key per intent so the same job is not created twice.
  • Separate dashboards: Z.ai points on one side, Sume spend on the other.

Reading the fine print on your plan

Quota rules differ by plan and can change. The GLM-5.3 page states the off-peak rate, but it does not define the peak window in the text read here, so check your plan page for the hours before you schedule anything around them.

Treat the discount as a nice-to-have, not as a design constraint. The render, the part with a dollar cost, runs when you submit it, and Sume's jobs_wait and receipts do not depend on the hour.

Sources

Related posts

More in Pricing

All Pricing posts

Written by Sume