Sume's GLM 5.3 Flash card is $0.075 in; Z.ai's page says $0.15

Sume's registry prices GLM 5.3 Flash at $0.075 input and $0.25 output per million. Z.ai's pricing page lists $0.15 and $0.50. Here is what to do with that gap.

4 min readSume
All posts

What differs

Two numbers for the same model name do not match as of 2026-10-08. Sume's registry has a rate card for GLM 5.3 Flash with input $0.075, output $0.25, cache read $0.015 and no cache write charge. The Z.ai pricing page lists GLM-5.3-Flash at $0.15 input, $0.03 cached input and $0.50 output per million tokens. Sume's figures are exactly half of Z.ai's.

I have not found anything in the repo or on the page that explains the gap.

Side by side

Z.ai also lists FlashX and the full GLM-5.3 on the same page. Sume lists none of those, and no row called Fast exists on the Z.ai page or in the registry.

GLM pricing, per million tokens (Sume agent model registry and rate cards on origin/main, read 2026-10-08; Z.ai page read 2026-10-08)
ModelInputCached inputOutputIn Sume registry
GLM-5.3-Flash (Z.ai)$0.15$0.03$0.50Row named GLM 5.3 Flash
GLM 5.3 Flash (Sume card)$0.075$0.015$0.25Yes
GLM-5.3-FlashX (Z.ai)$0.37$0.075$1.25No
GLM-5.3 (Z.ai)$1.40$0.26$4.40No
GLM 5.3 Fastnot on the Z.ai pageNo

Why it matters for a bill

On a turn of 30,000 input and 2,000 output tokens the Sume card gives 30,000 x 0.075 / 1M = $0.00225 plus 2,000 x 0.25 / 1M = $0.0005, so $0.00275. At the Z.ai figures it would be $0.0055. Either way the model-token cost is a fraction of a cent and is dwarfed by media: one second of Seedance 2 at 720p is $0.378 at Sume.

If you are budgeting against the lower card, check your receipt: the Format runs docs say usage.debited_usd_micros is the amount the wallet actually deducted, including the LLM row.

What I would do

Treat the receipt as the source of truth, not either price table. Run one small Format with model set to the GLM row, read debited_usd_micros, and size your wallet from that. If Z.ai's price is what you expect, tell Sume support the numbers disagree.

Questions to ask before trusting either number

A pricing mismatch between a vendor page and a reseller card has several ordinary explanations. A card may be copied from an aggregator that carries a different rate, a launch discount may be in force on one side, or the page may have changed since the card was written. The repo comment I could read for this card gives the rates and nothing about their source or date, and I did not find anything that settles it.

So the useful move is empirical. Run one Format with the GLM row, note the token counts if the receipt shows them, and divide the debited amount by what each price table predicts. Whichever table matches is the one Sume actually bills. Do that before you move a workload on the strength of 'half price'.

If you decide to move a workload, change one thing at a time. Keep the Format, input and cap the same, change only the model id, and compare debited_usd_micros for the two runs. A cheaper card only saves money if the cheaper model also finishes the task in a similar number of turns, and that is the part no price table can tell you. Keep a note of the date, since both the Z.ai page and Sume's card can change.

  • Sume card: $0.075 / $0.015 / $0.25 per million (input, cache read, output).
  • Z.ai page: $0.15 / $0.03 / $0.50 per million for GLM-5.3-Flash.

Sources

Related posts

More in Models

All Models posts

Written by Sume