Nightly Sume schedule: model choice and the 30-day cap math
A nightly Scheduled run caps generation at $1.00 by default, so 30 nights is at most $30. How the model's token cost sits beside that cap, with the arithmetic.

A nightly Sume schedule can spend at most the per-run generation cap times the number of nights. With the default cap of $1.00 per run, 30 nights is at most $30.00 of generation spend. The language model that reads your instructions is a second, separate cost, and it bills a different wallet.
Two costs, two wallets
A schedule is a saved Agents automation with instructions, a model, a cron expression, and a spend cap. When it fires, Sume runs it as an Agent in a fresh thread. The Scheduled docs say the cap is the maximum one run can spend on generation, and that the effective default is $1.00 per run if you set nothing.
The run receipt docs are blunt about the limit of that number: usage.billable_amount_usd_micros is generation spend only. It does not include the LLM turn of the agent, which bills the separate Agent wallet. So the cap is a ceiling on renders, music, and other media, not on the model that plans them.
The 30-night ceiling
The ceiling is linear, so you can choose it before you ship the schedule.
| Per-run cap | Nights | Maximum generation spend |
|---|---|---|
| $1.00 (default) | 30 | 30 x $1.00 = $30.00 |
| $0.40 | 30 | 30 x $0.40 = $12.00 |
| $2.50 | 30 | 30 x $2.50 = $75.00 |
Where the model token cost fits
Take Claude Haiku 5.5 as an example of a cheap planning model. Anthropic's page, read on 2026-10-08, lists $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100K tokens. Assume a nightly run that reads 60,000 tokens and writes 8,000 tokens in total.
Over 30 nights that is 1.8 million input tokens and 0.24 million output tokens. At list price: 1.8 x $0.10 = $0.18 for input, and 0.24 x $0.50 = $0.12 for output, so $0.30 for the month. These are Anthropic's list prices and an assumed token count. Sume's own Agent-wallet rate is a separate figure, so treat $0.30 as an order of magnitude, not a quote.
The comparison is the point. The planning model is cents per month in that example, while the generation cap is the number that can reach tens of dollars. Spend your attention on the cap.
Which models the repo lists
Sume's agent model catalog in the repository, read on 2026-10-08, has enabled rows for Haiku 5.5 and GLM 5.3 Flash, among others. Which rows a given environment shows depends on gates, so open the model picker in your own workspace before you plan around a name. It has no row for Mistral Large 4 and none for Cloudflare's or Amazon's open-weight decision models.
A checklist before you enable it
The points that matter here, in the order you will hit them:
- Set the cap on purpose. Do not rely on the $1.00 default if one render costs more than that.
- Remember an API caller can lower the cap for one run but never raise it.
- Use
on_active_runso a slow night does not stack runs; the default for schedules isskip. - Read
GET /v1/usagefor billing. The receipt figure is not an invoice.
After the first week
Do not trust the ceiling alone; measure against it. After seven nights, read the usage.billable_amount_usd_micros figure on each run receipt and compare it with generation_spend_cap_usd_micros. If every run lands near 20 percent of the cap, lower the cap so a runaway night cannot spend five times your normal. If runs hit the cap, find out whether the instructions ask for more media than you meant. Keep in mind that the receipt figure is a receipt, not an invoice; GET /v1/usage remains the authoritative billing record. Also check for skipped runs: they mean one night overran into the next, and the default overlap policy dropped the second trigger instead of stacking it.
Sources
Related posts
More in Agents
- One Sume webhook endpoint for three run types: verify, route
Format, action and agent runs share one signature scheme. Verify HMAC-SHA256 over timestamp.body, refuse an empty secret, then route on the event field.
- Pro key: 300 writes and 12,000 reads a minute for an agent on Sume MCP
On Sume Pro, each key gets 300 writes and 12,000 reads per minute. An agent polling 20 jobs every 5 s uses 240 reads, 2 percent of the read budget.
- Scheduled run spend cap: a per-call number can only lower it
A Sume schedule's default generation cap is $1.00. A per-run number is clamped to the lower of the request and the schedule cap, and 0 is rejected.
- A Sume image job at 50 seconds: slow, queued, or stuck?
Sume's docs say an image still running after a 50 s jobs_wait is usually stuck, not slow. Check queued vs processing and events first; never submit twice.
Written by Sume