Strands Decider 2B scores 167 of 231: not a spend guard

The Strands Decider 2B card reports 167 of 231 correct on JevBench public (72.3%). Fine for routing, not payment control. Why Sume's spend cap decides.

5 min readSume
All posts

Amazon's Strands Decider 2B card reports 167 correct out of 231 on JevBench public, an accuracy of 0.723. That means 64 wrong answers, about 27.7 percent, on that benchmark. A small decision model with those numbers is useful for triage and routing. It should not be the only thing between an agent and a paid render.

What the model card says

The Hugging Face card, read on 2026-10-08, describes an Apache-2.0 model built on Qwen/Qwen3.5-2B-Base with a LoRA adapter. It is described as a small, fast decision model for agentic AI that does model routing, tool selection, argument checking, triage, guardrails, and evaluations. It reports a Brier score of 0.348 and an expected calibration error of 0.050, and a primary evaluation context of 4,096 tokens.

Strands Decider 2B card figures (read 2026-10-08)
MetricValue
JevBench public, correct167 of 231
Accuracy0.723
Wrong answers64 of 231 (27.7%)
Brier score0.348
ECE0.050
Primary evaluation context4,096 tokens

Why 72.3 percent is not a guard

The benchmark measures the model on a public question set, not on your questions. On your own traffic the rate can be better or worse, and you will not know which without labeled examples. Even at the benchmark rate, a gate that approves "is this render request fine?" would be wrong for roughly 277 of every 1,000 decisions (1,000 x 0.277).

A guardrail needs a hard limit, not a probability. Sume has hard limits at the API. Every Agent Completion must send generation_spend_cap_usd, and the docs say the cap replaces the interactive spend-approval prompt that a chat user gets. On hosted MCP, paid tools need an idempotency_key, and max_spend_usd is enforced when you provide it.

Where a small decision model does help

The points that matter here, in the order you will hit them:

  • Choosing which of a handful of tools to call first, where a wrong pick costs a retry.
  • Triaging incoming requests into a queue, with a human review lane for low-confidence answers.
  • Flagging a prompt for review before it reaches a paid path, as one layer among several.

The layering that works

Put the model first, the cap last. Use the decider's probability to decide whether to bother asking the agent at all, and set the cap to the most you would accept if the decider were wrong every time. Sume's agent model catalog in the repository has no Strands Decider row, so any such model runs in your own code.

What calibration buys you

The card's ECE of 0.050 says that, on that benchmark, the model's stated confidence sits within about five points of its observed accuracy on average. That is the useful property. It lets you route by confidence: act on high-confidence answers and send the rest to a person or to a stricter path. Choose the threshold from your own labeled sample, not from the card. For example, you might send every answer under 0.80 to review and measure how many reviews that creates.

Finally, remember the context. The card's primary evaluation used a 4,096-token context. A state block longer than that is outside what the numbers describe, so shorten the state before you trust the figures.

Sources

Related posts

More in Models

All Models posts

Written by Sume