On 1st October 2026, Amazon Web Services released Strands Decider 2B, an open-source model that picks from fixed answers and scores how sure it is. On a single Nvidia RTX 3090, it returns a decision in a median of about 115 milliseconds.
The Brief
- Amazon Web Services built Strands Decider 2B on the Qwen3.5-2B language model and removed its ability to generate text.
- Strands Decider 2B placed 3rd of 33 models in its size class on JevBench for accuracy and calibration combined.
- Developers can download the weights, training data and training scripts from GitHub and Hugging Face and run everything locally.
- OpenAI announced a similar offering the same week, as bigger AI players follow the design TypeSafe introduced with Jev.
Marc Brooker, a distinguished engineer at Amazon (NASDAQ: AMZN), started the project after TypeSafe launched Jev in September. His homebrew version briefly topped the Jevbench ranking for its size. AWS engineers then cleaned it up and shipped it through Strands Labs, the company’s group for agent tools.
Brooker told TechCrunch the push came from AWS customers. Their agent workflows didn’t always need the capability, or the cost, of a full large language model (LLM). For Amazon’s cloud business, the release is a pitch for cheaper, faster steps inside agents customers already run.
🆕 AWS added Strands Decider 2B to strands-labs, a small decision model for agentic AI optimized for fast experimentation and local development. It runs locally, returns answers in under 100 milliseconds, and is fully open source with all training data and scripts included.… pic.twitter.com/LzbtjKZoeM
— AWS Newsroom (@AWSNewsroom) October 1, 2026
Amazon swapped the text head for a pointer
Strands Labs took the torso of Qwen3.5-2B and cut off the head that turns its output into words. A pointer head of just over 1 million parameters now scores each answer option instead. A rank-16 LoRA adapter (a lightweight fine-tuning layer) handles the rest. The public release is v19; an earlier slot-head design performed significantly worse, the team said.
The trade-off is blunt. Because it only answers from the options it’s handed, the model runs fast and never goes off script. AWS says each decision also carries a reliability score that frontier LLM inference APIs don’t expose. The catch is a single parallel pass, which leaves it significantly worse than reasoning models on complex problems. Coding, chatbots and document summaries are off the table.
Excluding models just over 2 billion parameters, it ranks first of 30 on JevBench’s public set. AWS measured calibration with the Brier score, a check on whether stated confidence matches real accuracy. The model also got 100% of JevBench’s easy tasks right. On an M3 MacBook, small tasks took a median of about 153 milliseconds, and latency climbed roughly with task size. One caveat: the published latency chart measured v18, one version behind the release.
A cheap check before an agent acts
The clearest use case in AWS’s own announcement sits right before a tool call. A deliberately eager demo agent gets asked “What’s the weather?” and guesses a city anyway. Before get_weather runs, Strands Decider answers two yes/no questions about grounded arguments and premature calls. The agent then asks which city the user meant.
That check plugs into Strands’ intervention system, where a before_tool_call handler can return Proceed, Deny, Confirm or Guide. AWS said it picked the questions, threshold and policy by hand, and called the demo an illustration only. Its case is that “a decision this cheap can sit in a path where an LLM call never could.”
Teams running AI agents can install it with pip install strands-decider and load the StrandsAgents/strands-decider-2B-hobson-v19 model. A gate like this helps reduce the risk of agents acting on invented arguments. Hand-set thresholds still need testing on real traffic first. Two questions stay open. How well does it hold up outside JevBench’s public set, and how far can AWS lower its latency floor?
The Bottom Line
TypeSafe founder and CEO Diogo Almeida isn’t worried. He said rivals “might be underestimating the difficulty of making the models actually smart,” and sees no real competition yet. Brooker doesn’t expect frontier labs to own the category either, since interesting niche models cost hundreds or thousands of dollars to build. The hard part, he said, is pushing accuracy and calibration without eroding language skills and general knowledge.
TypeSafe named Jev after economist William Stanley Jevons, who argued that a cheaper resource can end up in higher demand. Amazon’s bet follows that logic, and the Strands team says libraries for wiring decision models into agents will land in its GitHub repo soon.