Skip to content

Strands Decider 2B makes agent choices in 115 ms

AWS's Strands Decider 2B is an open-source model that picks from a list instead of writing text. It answers in 115 ms on an RTX 3090 and scores 72.3% on JevBench.

By Tech AI Wire Team

4 min read

XLinkedIn
Screenshot of the Strands Agents blog post Introducing Strands Decider 2B: a small, open source, decision model, dated October 1, 2026.

By the numbers

median latency on an RTX 3090
115 ms
accuracy on JevBench v1's 231 tasks
72.3%
parameters in the pointer head that replaces text output
1M
license, per Tech Times
Apache 2.0

AWS's Strands Labs team released Strands Decider 2B on October 1, 2026, an open-source model that makes decisions for AI agents without writing any text. Given a question and a list of options, it scores each option and returns its pick with a confidence value. Strands' announcement puts its median response time at 115 milliseconds on an RTX 3090 graphics card. That makes it cheap enough to run before every expensive call to a large language model.

What a decision model does

Most AI agents ask a large language model, or LLM, to make small choices all day. Which tool should run next? Which model should answer? Is this request allowed? The LLM writes its answer as text, token by token, which is slow and costs money.

A decision model skips the writing. It reads the question and the list of allowed answers, then gives each answer a score. The options are supplied at the moment you ask, so the same model can choose between tools in one call and between policies in the next.

Strands lists these jobs for it:

  • model routing, which sends each request to the right model
  • tool selection
  • evaluations
  • guardrails, which block requests that break a policy
  • memory and context management

The limits are just as clear. AWS says the model "cannot write code, summarize documents, carry on a conversation, or excel at complex reasoning tasks," Techstrong.ai reports.

How it is built

Decider 2B starts from Alibaba's Qwen3.5-2B-Base model. Strands removed the language modeling head, the final layer that turns the model's internal state into words. In its place sits a "pointer head" of about 1 million parameters, according to SiliconANGLE.

Tech Times explains how the head works. It compares the model's reading of the input with each candidate option, using a dot product, a simple similarity score. The result is a ranked list of options with probabilities. Strands' own post says the model was adapted with a rank-16 LoRA adapter, a small add-on trained instead of the full set of weights.

Reports give the size as about 1.9 billion parameters. Strands rounds it to 2 billion in the name.

The numbers

MeasureResultReported by
Median latency, RTX 3090115 msStrands, Tech Times
95th-percentile latency299 msTech Times
Latency, Apple M3 MacBook153 msStrands
JevBench v1 accuracy72.3% (167 of 231 tasks)Tech Times, Techstrong.ai
Easy-tier tasks100%Techstrong.ai
Hard-tier tasks50.5%Techstrong.ai
Expected calibration error0.052Tech Times

JevBench is a public test set for decision models. Strands says Decider 2B placed third of 33 models in the 2B class. Leaving out models just over 2 billion parameters, it ranks first of 30.

Calibration error measures how well a model's confidence matches reality. A low number like 0.052 means that when the model says it is 90% sure, it is right close to that often. That matters more here than raw accuracy, as the next section explains.

The model and its code are free to download. Tech Times reports the Apache 2.0 license. The weights sit on Hugging Face under the StrandsAgents organization, and the training code is on GitHub at strands-labs/strands-decider.

A growing class of models

Decision models are becoming their own category. Techstrong.ai says Strands credits TypeSafe AI's Jev, released in September, with drawing attention to the approach. An open-weights rival, Laya, answered Jev the same month. Cloudflare followed on October 4 with Clef, open-weight models that return decisions.

AWS frames the trade-off in speed. The model "eliminates the need to generate text, which eats up vast amounts of tokens and increases response latency," it told SiliconANGLE.

What this means for developers

  • Put it in front of your LLM, not in place of it. Use Decider 2B for the fixed-choice steps in an agent: routing, tool picks and policy checks. Leave writing and reasoning to the large model.
  • Use the confidence score as a gate. Because the model is well calibrated, you can set a threshold. Above it, act on the pick. Below it, hand the decision to a larger model or a person.
  • Keep it away from hard calls. A 50.5% score on hard tasks is close to a coin flip on two options. Test your own decisions and sort them by difficulty before you trust it.
  • Write options the model can tell apart. The pointer head compares the input with each option's text. Clear, distinct option labels give it more to work with than near-duplicates.
  • Run it locally. At about 2 billion parameters, it fits on a consumer GPU or a laptop. That keeps routine decisions off the meter and off the network.

Build a small test set from your agent's real logs first. Then compare Decider 2B's picks with what your LLM chose, and measure how often they agree.

Sources

  1. Introducing Strands Decider 2B: a small, open source, decision model - Strands Agents
  2. AWS debuts Strands Decider 2B, a first lightweight decision model for accelerate agentic workflows - SiliconANGLE
  3. AWS Releases Decision Model for AI Agents That Routes Without Generating Any Text - Tech Times
  4. AWS Explores Decision Models With Strands Decider 2B - Techstrong.ai

Related articles

The weekly digest

One email every Friday with the week's top stories from all six desks: AI, dev tools, coding, the tech industry, startups and what's next. Free, no spam.

Unsubscribe anytime with one click.