Strands Decider 2B, a Model That Writes No Text, Ships Under Apache 2.0 - One Decision in 115 ms Median
On October 1, 2026, strands-labs released Strands Decider 2B, a 1.9-billion-parameter decision model with its text-generation ability removed, under Apache 2.0. It answers only in three forms - pick, yes/no, or score - and returns in a median of 115 ms on an RTX 3090. Weights, training data, and training scripts are all public.
On October 1, 2026, strands-labs — the experimental arm of Strands Agents, a family of open-source agent-building projects — released Strands Decider 2B. It is a model whose ability to write text has been stripped out: the language-modelling head of a pretrained LLM is gone1. What remains answers in exactly three forms — choose one item from a list (choice), answer yes or no (noul), or place something on an ordered scale (score). In return, a single question comes back in a median of 115 ms on an NVIDIA RTX 3090, or 153 ms on an Apple-silicon M3 Pro for tasks under 300 tokens2.
The repository puts the parameter count at 1.9 billion2. Weights sit on Hugging Face; the code, the training data, and the training scripts sit on GitHub, all under the Apache License 2.0. A pip install strands-decider gets you the CLI, and the model will serve from a CPU or a GPU on your own machine1.
What Removing the Language-Modelling Head Buys
Inside, the torso of Qwen3.5-2B-Base keeps its weights while the LM head is swapped for a pointer head of roughly a million parameters. To rank the candidates, that head compares two hidden states: the one sitting at the <answer> position, and the one at each candidate’s last token. There is a single forward pass and nothing after it — no sampling, no decode loop. Adaptation on the torso side is handled by a rank-16 LoRA adapter2. Rather than bolting a new role onto an existing model through fine-tuning, what has been rebuilt here is arguably the exit path itself.
That choice has knock-on effects. Since the pointer head carries no parameters per option, it has nowhere to encode a bias such as “the first option tends to be right,” a question can carry any number of options, and the label set is decided by the caller rather than frozen into the weights. The small parameter count may matter less in practice than that last property.
Confidence is the other difference. strands-labs points out that a decision model attaches a reliability score to every answer, and says frontier LLM inference APIs offer no equivalent. The repository records that on unseen short classification tasks, answers the model rates at 0.9 confidence or higher prove correct in roughly 95% of cases, and that anything beneath that line should be verified or escalated to a person2. Whether you can draw that line at a threshold is what changes in guardrail design.
Posing several questions about one input is also cheap by construction: the input gets read a single time, and every extra question costs only the tokens it brings with it. The CLI lets you stack --choice, --noul, and --score in a single invocation.
What It Is Being Offered For
The use cases named are the decisions that happen inside an agent: model routing, tool selection, guardrails, evaluations and the like. On top of that sits what the announcement calls a hybrid agent — leave the hard calls to an LLM and hand the rote ones to a decision model, cutting both cost and latency1.
If the branch points inside your agent are currently plain LLM calls, several of them are candidates for replacement. On model routing, the recent movement has included cloud-side offerings — Ramp opened its in-house LLM router to outside customers — whereas this one drops 1.9 billion parameters onto a GPU already sitting on your desk. With the round trip to a cloud API gone, there is room for both latency and billing to move.
A worked example lives in the repository at examples/strands/, with the model embedded in a Strands agent. An intervention on before_tool_call gates a weather tool so the agent stops calling it on a city the user never named. Two yes/no questions do the gating: whether each argument value has a basis in what the user genuinely stated (args_grounded), and whether firing the tool now, ahead of any clarification, is premature (premature). The intervention returns one of four typed actions — Proceed, Deny, Confirm, or Guide — where Confirm stops and asks a person, and Guide hands the turn back to the model with feedback instead of blocking the call1. In this example both the agent and the decision model run locally, with Amazon Bedrock supplying the default LLM. The repository has also adopted Amazon’s open-source code of conduct2.
strands-labs states plainly that the example is an illustration and not a recommendation: every question, the cutoff and the policy were picked by hand. Libraries for integrating decision models are still being built, and readers are pointed at the repository for updates. For now, writing the question text and picking the threshold stays on the user’s side.
Where It Does Not Fit, and How to Read the Benchmark
The publishers are blunt about the limits. Because every output is produced in one parallel pass, the model is far weaker than a reasoning model at hard problems, and with no ability to write text, coding, chatbots and document summarisation — the everyday LLM jobs — are out of scope1. The Hugging Face model card lists two further limits: weakness on long, multi-step documents, and calibration that was fitted only on short classification tasks3.
Accuracy is reported on JevBench, which the repository calls an outside benchmark built for models of this kind: 0.723 on the 231-task public set (167 correct), a Brier score of 0.342 and an expected calibration error of 0.0522. Those figures were taken at a 3072-token window, which is what the run had been preregistered for; widen it to the 4096-token window the recipe saves and v19 reaches 168. Split into tiers by the repository, the easy tier comes in at 1.000, standard at 0.875, and hard at 0.505. strands-labs characterises its accuracy and calibration as “3rd of 33 in the 2B class, and 1st of 30 excluding the just-over-2B models”1.
There is a caveat about reading those figures at all. The set holds only 231 tasks; retraining the v17 recipe six times gave a spread of 3.2 tasks, and the guidance is therefore to call any gap under roughly ten tasks between two lone runs undecided. That caveat has teeth: the published reference model is still v19, because the newer v20 scored 169 correct out of 231 with a 4096-token window, one above v19’s 168, and that margin was judged to fall within retrain noise2.
As for the category, strands-labs frames decision models as a class that has drawn attention since TypeSafe AI launched Jev, and casts this release as its own contribution to it. Keeping the model small enough to retrain on equipment you already own is also described as a way to encourage experimentation: one RTX 3090 with 24 GiB takes roughly 11 hours for the full recipe, while eight H100s with FAST=1 finish it in 70 minutes2. Training, though, presumes NVIDIA GPUs on a host running Linux or WSL2, so unlike inference it will not finish on a Mac.
Sources
- Introducing Strands Decider 2B: a small, open source, decision model - Strands Agents official blog (October 1, 2026)
- strands-labs/strands-decider - Official repository (README, license, evaluation results)
- StrandsAgents/strands-decider-2B-hobson-v19 - Hugging Face model card
Was this article helpful?
Thank you!
Received. Thank you!