Strands Decider 2B scores 167 of 231: not a spend guard
The Strands Decider 2B card reports 167 of 231 correct on JevBench public (72.3%). Fine for routing, not payment control. Why Sume's spend cap decides.

Amazon's Strands Decider 2B card reports 167 correct out of 231 on JevBench public, an accuracy of 0.723. That means 64 wrong answers, about 27.7 percent, on that benchmark. A small decision model with those numbers is useful for triage and routing. It should not be the only thing between an agent and a paid render.
What the model card says
The Hugging Face card, read on 2026-10-08, describes an Apache-2.0 model built on Qwen/Qwen3.5-2B-Base with a LoRA adapter. It is described as a small, fast decision model for agentic AI that does model routing, tool selection, argument checking, triage, guardrails, and evaluations. It reports a Brier score of 0.348 and an expected calibration error of 0.050, and a primary evaluation context of 4,096 tokens.
| Metric | Value |
|---|---|
| JevBench public, correct | 167 of 231 |
| Accuracy | 0.723 |
| Wrong answers | 64 of 231 (27.7%) |
| Brier score | 0.348 |
| ECE | 0.050 |
| Primary evaluation context | 4,096 tokens |
Why 72.3 percent is not a guard
The benchmark measures the model on a public question set, not on your questions. On your own traffic the rate can be better or worse, and you will not know which without labeled examples. Even at the benchmark rate, a gate that approves "is this render request fine?" would be wrong for roughly 277 of every 1,000 decisions (1,000 x 0.277).
A guardrail needs a hard limit, not a probability. Sume has hard limits at the API. Every Agent Completion must send generation_spend_cap_usd, and the docs say the cap replaces the interactive spend-approval prompt that a chat user gets. On hosted MCP, paid tools need an idempotency_key, and max_spend_usd is enforced when you provide it.
Where a small decision model does help
The points that matter here, in the order you will hit them:
- Choosing which of a handful of tools to call first, where a wrong pick costs a retry.
- Triaging incoming requests into a queue, with a human review lane for low-confidence answers.
- Flagging a prompt for review before it reaches a paid path, as one layer among several.
The layering that works
Put the model first, the cap last. Use the decider's probability to decide whether to bother asking the agent at all, and set the cap to the most you would accept if the decider were wrong every time. Sume's agent model catalog in the repository has no Strands Decider row, so any such model runs in your own code.
What calibration buys you
The card's ECE of 0.050 says that, on that benchmark, the model's stated confidence sits within about five points of its observed accuracy on average. That is the useful property. It lets you route by confidence: act on high-confidence answers and send the rest to a person or to a stricter path. Choose the threshold from your own labeled sample, not from the card. For example, you might send every answer under 0.80 to review and measure how many reviews that creates.
Finally, remember the context. The card's primary evaluation used a 4,096-token context. A state block longer than that is outside what the numbers describe, so shorten the state before you trust the figures.
Sources
Related posts
More in Models
- sume/music-auto or pinned lyria-3.5: what changes in the job
Music Router takes sume/music-auto (Lyria 3.5 today), lyria-3.5 or lyria-3-pro. What job.model and routed_model echo back, and why the price stays the same.
- Swap four people in one video: H3 Max Recast bills per second
h3-max-recast takes one source video and 1 to 4 person photos, 5 to 30 seconds. Price is $0.375 or $0.5625 a second whether you swap one person or four.
- Hy Image 3.5 Preview is not in Sume's catalog: what to call instead
Sume's image catalog does not list Hy Image 3.5 Preview. Check with one jq call, then pick a listed model by price: Soul, gpt-image-2.5, Seedream or Flux.
- Text-in-image bake-off: 20 prompts on 4 Sume models costs $7.00
Test text rendering in AI images for $7.00: 20 prompts on Ideogram 4.5, GPT Image 2.5, Qwen Image Max and Nano Banana 2.1, with the cost of each on Sume.
Written by Sume