Runway Model Router credit ceiling vs Sume max_spend_usd
Runway's per-modality credit ceiling removes costly models before ranking and errors if none fit. Sume's max_spend_usd caps one call. How the two caps differ.

A Runway Model Router config can carry a credit ceiling for each modality, and any model whose estimated cost for a request is above that ceiling is dropped before the cost, latency, or quality preference is applied. If nothing fits, the request is rejected with an explicit error. Sume has no router-level ceiling: spend is capped per call with the optional max_spend_usd on MCP paid tools, and a request the balance cannot cover fails with 402 insufficient_credits.
Runway's side is from Configuring a Model Router and the API changelog (Model Router entry, July 23, 2026), read on 2026-10-03. Sume's side is from MCP tools and gates and errors and credits.
How does Runway's ceiling work?
The configuration page describes the ceiling as per modality, so video, image, and audio each get their own number. Order matters: models over the ceiling are excluded first, then the router applies your preference among what is left (cost, latency, or quality).
The same page says the config ID is set once at creation and cannot change, while name, description, and preferences stay editable and apply only to later requests. The changelog adds allow and deny lists and a dryRun preview.
| Control | Effect |
|---|---|
| Allow list | Only listed models are eligible; new models are not added |
| Deny list | Everything except listed models is eligible; new models join automatically |
| Credit ceiling per modality | Models estimated above it are excluded before ranking |
| Preference | Cost, latency, or quality, applied to what remains |
What is the Sume equivalent?
Sume's caps work at the call, not at a saved routing config. On hosted MCP, max_spend_usd is optional and enforced only when you provide it, and dry_run=true returns an admission and cost preview without submitting. Over the API, a job is admitted only when the balance can be reserved: reservation happens at submit, at provider list price times 1.25 for video, and a short balance returns 402 insufficient_credits before provider work starts.
With model: "sume/auto", Sume picks the family and does not disclose which one ran, so there is no ceiling you configure for Auto beyond the per-call cap and your balance.
Which cap fits which job?
Use a ceiling when many callers share one router and you want a rule that does not depend on each caller remembering a number. Use a per-call cap when each request has its own budget, as in an agent that has been told to spend at most two dollars on a clip.
The honest gap: Sume does not offer a saved config that excludes expensive models for every caller. If you want that, enforce it in your own gateway before you call Sume, by pinning the model ids you allow.
How do you enforce a ceiling yourself on Sume?
Keep an allowlist of catalog ids with a maximum duration and resolution for each, reject anything outside it in your code, and send max_spend_usd on MCP calls. For the API, read each model's limits from GET /v1/videos/models, which the video docs say differ per model, then compute the worst case before you submit.
Queue pressure is a separate control. Sume's generation admission page lists concurrency, queue capacity, submit rate limits, and balance as four different controls, so a full queue (429 queue_full) is not a spend problem.
What does 'no eligible model' teach you?
Runway's generating page says that when no model meets the configuration, the error names which constraints eliminated the options, with the usual fixes being to enable more models or raise the ceiling for that modality. That is a useful design: the failure is about policy, not about a provider being down.
Sume's closest analogue is the pair of submit-time errors. 402 insufficient_credits means the balance cannot cover the reserve; 429 queue_full means workspace concurrency and queue capacity are used up; 429 rate_limited means too many requests in the window. Each one tells you a different thing to fix, and none of them is a model-eligibility error.
What should an agent do when a cap blocks it?
Do not loop. A blocked request is a decision for a person or a policy: raise the cap, lower the quality tier, or shorten the clip. Give your agent the three options in its instructions and tell it to stop and report if none applies. On Sume, max_spend_usd is enforced only when provided, so an agent that omits it has no per-call cap beyond the balance.
Make the cap part of the prompt template, not something each caller remembers.
Sources
Related posts
More in Comparisons
- Runway Model Router dryRun: preview the pick, and what Sume Auto shows
Runway's Model Router has a dryRun preview, per-modality credit ceilings and cost, latency or quality goals. Sume's sume/auto never says which family ran.
- Runway Seedance 2.0 4K at 150 credits a second: cost
Runway bills Seedance 2.0 4K at 150 credits a second, $1.50 at its $0.01 credit. What a clip costs against 1080p, and what Sume's seedance-2 offers instead.
- Runway Seedance 2.5 80-credit minimum: when it applies
Runway's Seedance 2.5 has an 80-credit minimum, but a 4-second 480p clip already costs 80. When the floor matters, the 1080p math, and Sume's reserve.
- Schedule a weekly AI video: Sume Scheduled, Hermes cron or API
Three ways to run an AI video on a weekly clock with Sume: a dashboard schedule, a Hermes cron job, or plain cron calling a Format. What each can and cannot do.
Written by Sume