Strands Decider 2B makes agent choices in 115 ms
AWS's Strands Decider 2B is an open-source model that picks from a list instead of writing text. It answers in 115 ms on an RTX 3090 and scores 72.3% on JevBench.
4 min read

By the numbers
- median latency on an RTX 3090
- 115 ms
- accuracy on JevBench v1's 231 tasks
- 72.3%
- parameters in the pointer head that replaces text output
- 1M
- license, per Tech Times
- Apache 2.0
AWS's Strands Labs team released Strands Decider 2B on October 1, 2026, an open-source model that makes decisions for AI agents without writing any text. Given a question and a list of options, it scores each option and returns its pick with a confidence value. Strands' announcement puts its median response time at 115 milliseconds on an RTX 3090 graphics card. That makes it cheap enough to run before every expensive call to a large language model.
What a decision model does
Most AI agents ask a large language model, or LLM, to make small choices all day. Which tool should run next? Which model should answer? Is this request allowed? The LLM writes its answer as text, token by token, which is slow and costs money.
A decision model skips the writing. It reads the question and the list of allowed answers, then gives each answer a score. The options are supplied at the moment you ask, so the same model can choose between tools in one call and between policies in the next.
Strands lists these jobs for it:
- model routing, which sends each request to the right model
- tool selection
- evaluations
- guardrails, which block requests that break a policy
- memory and context management
The limits are just as clear. AWS says the model "cannot write code, summarize documents, carry on a conversation, or excel at complex reasoning tasks," Techstrong.ai reports.
How it is built
Decider 2B starts from Alibaba's Qwen3.5-2B-Base model. Strands removed the language modeling head, the final layer that turns the model's internal state into words. In its place sits a "pointer head" of about 1 million parameters, according to SiliconANGLE.
Tech Times explains how the head works. It compares the model's reading of the input with each candidate option, using a dot product, a simple similarity score. The result is a ranked list of options with probabilities. Strands' own post says the model was adapted with a rank-16 LoRA adapter, a small add-on trained instead of the full set of weights.
Reports give the size as about 1.9 billion parameters. Strands rounds it to 2 billion in the name.
The numbers
| Measure | Result | Reported by |
|---|---|---|
| Median latency, RTX 3090 | 115 ms | Strands, Tech Times |
| 95th-percentile latency | 299 ms | Tech Times |
| Latency, Apple M3 MacBook | 153 ms | Strands |
| JevBench v1 accuracy | 72.3% (167 of 231 tasks) | Tech Times, Techstrong.ai |
| Easy-tier tasks | 100% | Techstrong.ai |
| Hard-tier tasks | 50.5% | Techstrong.ai |
| Expected calibration error | 0.052 | Tech Times |
JevBench is a public test set for decision models. Strands says Decider 2B placed third of 33 models in the 2B class. Leaving out models just over 2 billion parameters, it ranks first of 30.
Calibration error measures how well a model's confidence matches reality. A low number like 0.052 means that when the model says it is 90% sure, it is right close to that often. That matters more here than raw accuracy, as the next section explains.
The model and its code are free to download. Tech Times reports the Apache 2.0 license. The weights sit on Hugging Face under the StrandsAgents organization, and the training code is on GitHub at strands-labs/strands-decider.
A growing class of models
Decision models are becoming their own category. Techstrong.ai says Strands credits TypeSafe AI's Jev, released in September, with drawing attention to the approach. An open-weights rival, Laya, answered Jev the same month. Cloudflare followed on October 4 with Clef, open-weight models that return decisions.
AWS frames the trade-off in speed. The model "eliminates the need to generate text, which eats up vast amounts of tokens and increases response latency," it told SiliconANGLE.
What this means for developers
- Put it in front of your LLM, not in place of it. Use Decider 2B for the fixed-choice steps in an agent: routing, tool picks and policy checks. Leave writing and reasoning to the large model.
- Use the confidence score as a gate. Because the model is well calibrated, you can set a threshold. Above it, act on the pick. Below it, hand the decision to a larger model or a person.
- Keep it away from hard calls. A 50.5% score on hard tasks is close to a coin flip on two options. Test your own decisions and sort them by difficulty before you trust it.
- Write options the model can tell apart. The pointer head compares the input with each option's text. Clear, distinct option labels give it more to work with than near-duplicates.
- Run it locally. At about 2 billion parameters, it fits on a consumer GPU or a laptop. That keeps routine decisions off the meter and off the network.
Build a small test set from your agent's real logs first. Then compare Decider 2B's picks with what your LLM chose, and measure how often they agree.
Sources
- Introducing Strands Decider 2B: a small, open source, decision model - Strands Agents
- AWS debuts Strands Decider 2B, a first lightweight decision model for accelerate agentic workflows - SiliconANGLE
- AWS Releases Decision Model for AI Agents That Routes Without Generating Any Text - Tech Times
- AWS Explores Decision Models With Strands Decider 2B - Techstrong.ai
Related articles

Cloudflare Clef: open-weight models that return decisions
Cloudflare's Clef (27B) and Clef-flash (9B) answer typed questions instead of writing text, at 209 ms and 39 ms median latency, under Apache 2.0.

Qwen3.5-9B spends 32KB per token on KV cache
Qwen3.5-9B runs full attention on only 8 of its 32 layers, so its KV cache costs 32KB per token. That ratio decides whether it fits in 8GB.

EmbeddingGemma 2 puts multimodal search on a phone
Google DeepMind's EmbeddingGemma 2 maps text, code, images, video and audio into one vector space with 740M parameters. Text-only use needs about 191 MB of RAM.
The weekly digest
One email every Friday with the week's top stories from all six desks: AI, dev tools, coding, the tech industry, startups and what's next. Free, no spam.