Nightly Sume schedule: model choice and the 30-day cap math

A nightly Scheduled run caps generation at $1.00 by default, so 30 nights is at most $30. How the model's token cost sits beside that cap, with the arithmetic.

5 min readSume
All posts

A nightly Sume schedule can spend at most the per-run generation cap times the number of nights. With the default cap of $1.00 per run, 30 nights is at most $30.00 of generation spend. The language model that reads your instructions is a second, separate cost, and it bills a different wallet.

Two costs, two wallets

A schedule is a saved Agents automation with instructions, a model, a cron expression, and a spend cap. When it fires, Sume runs it as an Agent in a fresh thread. The Scheduled docs say the cap is the maximum one run can spend on generation, and that the effective default is $1.00 per run if you set nothing.

The run receipt docs are blunt about the limit of that number: usage.billable_amount_usd_micros is generation spend only. It does not include the LLM turn of the agent, which bills the separate Agent wallet. So the cap is a ceiling on renders, music, and other media, not on the model that plans them.

The 30-night ceiling

The ceiling is linear, so you can choose it before you ship the schedule.

Generation ceiling per schedule, from the Create a schedule page (read 2026-10-08)
Per-run capNightsMaximum generation spend
$1.00 (default)3030 x $1.00 = $30.00
$0.403030 x $0.40 = $12.00
$2.503030 x $2.50 = $75.00

Where the model token cost fits

Take Claude Haiku 5.5 as an example of a cheap planning model. Anthropic's page, read on 2026-10-08, lists $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100K tokens. Assume a nightly run that reads 60,000 tokens and writes 8,000 tokens in total.

Over 30 nights that is 1.8 million input tokens and 0.24 million output tokens. At list price: 1.8 x $0.10 = $0.18 for input, and 0.24 x $0.50 = $0.12 for output, so $0.30 for the month. These are Anthropic's list prices and an assumed token count. Sume's own Agent-wallet rate is a separate figure, so treat $0.30 as an order of magnitude, not a quote.

The comparison is the point. The planning model is cents per month in that example, while the generation cap is the number that can reach tens of dollars. Spend your attention on the cap.

Which models the repo lists

Sume's agent model catalog in the repository, read on 2026-10-08, has enabled rows for Haiku 5.5 and GLM 5.3 Flash, among others. Which rows a given environment shows depends on gates, so open the model picker in your own workspace before you plan around a name. It has no row for Mistral Large 4 and none for Cloudflare's or Amazon's open-weight decision models.

A checklist before you enable it

The points that matter here, in the order you will hit them:

  • Set the cap on purpose. Do not rely on the $1.00 default if one render costs more than that.
  • Remember an API caller can lower the cap for one run but never raise it.
  • Use on_active_run so a slow night does not stack runs; the default for schedules is skip.
  • Read GET /v1/usage for billing. The receipt figure is not an invoice.

After the first week

Do not trust the ceiling alone; measure against it. After seven nights, read the usage.billable_amount_usd_micros figure on each run receipt and compare it with generation_spend_cap_usd_micros. If every run lands near 20 percent of the cap, lower the cap so a runaway night cannot spend five times your normal. If runs hit the cap, find out whether the instructions ask for more media than you meant. Keep in mind that the receipt figure is a receipt, not an invoice; GET /v1/usage remains the authoritative billing record. Also check for skipped runs: they mean one night overran into the next, and the default overlap policy dropped the second trigger instead of stacking it.

Sources

Related posts

More in Agents

All Agents posts

Written by Sume