Franklin AI Explainer

Strands releases Decider 2B for fast local choices inside agent workflows

Key Takeaways

  • The open-source decision model chooses among supplied options rather than generating text.
  • Strands reports local latency and calibration results and publishes training artifacts, w The open-source decision model chooses among supplied options rather than generating text.
  • Strands reports local latency and calibration results and publishes training artifacts, while warning that this design is not suited to chat, coding, or complex reasoning.
  • Strands Decider 2B makes choices from a supplied set of options instead of composing an open-ended answer.
  • The October 1 release introduces a small open-source decision model aimed at local experimentation and the repeated, constrained decisions inside agent workflows.

Strands Decider 2B makes choices from a supplied set of options instead of composing an open-ended answer. The October 1 release introduces a small open-source decision model aimed at local experimentation and the repeated, constrained decisions inside agent workflows.

The team says it has released code, model weights, training data, and training scripts. Its examples include choosing a support team, classifying policy questions, and checking a proposed tool call before an agent executes it.

A model built to select, not write

The architecture starts with the pretrained Qwen3.5-2B model body, removes its language-model head, and replaces it with a pointer head. That head scores supplied answer options using their hidden states and the answer-position representation. The body is fine-tuned with a rank-sixteen LoRA adapter.

The result cannot generate ordinary text. It can choose from a defined set or assign simple scores, but the publisher explicitly says the design is unsuitable for coding, chatbots, document summarization, and other generative tasks. It also performs worse on complex problems than reasoning models.

That constraint can be useful when a workflow needs a bounded answer such as a destination team or a yes/no judgment. Producing a valid option still does not establish that the selected option is correct.

Accuracy, calibration, and latency have separate evidence

The team evaluates accuracy on JevBench's public set and calibration through Brier score on that same set. It reports third place among thirty-three models in the 2B class, or first among thirty when just-over-2B entries are excluded. Those rankings use the publisher's stated comparison group, not an independently established lead across every small model.

For local inference, the announcement reports median decision latency around 115 milliseconds on an Nvidia RTX 3090 and around 153 milliseconds for small tasks on an M3 MacBook. Task size affects latency. The displayed RTX 3090 latency plot measures version eighteen, while the released model is version nineteen, so the graph should not be treated as an exact measurement of every request on the current release.

The article's broader reference to tens-of-milliseconds answers likewise does not override these hardware- and workload-specific medians. Developers would need to measure their own task sizes and deployment conditions.

Checking an agent before a tool runs

A supplied example connects a local Strands agent and local decision model with a default language model from Amazon Bedrock. The agent is intentionally eager: asked for the weather without a location, it guesses a city and attempts a tool call.

Before execution, the decision model evaluates whether the arguments came from the user and whether clarification is still required. A small policy then guides the agent back to ask for a city. The intervention system can proceed, deny, request confirmation, or hand feedback back to the model.

This is an illustration with hand-chosen questions, thresholds, and policy, not a recommended universal guardrail. The local decision component also does not make the entire demonstration local-only, because its main language model uses Bedrock.

The release offers a concrete building block for routing and checks that do not need free-form generation. Its limited answer space, visible training artifacts, and local execution make experiments accessible. The CLI accepts explicit choices, so developers can test their own bounded questions without first building a complete agent. Whether it improves a production agent depends on the questions, policy, and error costs around those decisions, rather than the model's size alone.

Our read

Franklin AI Take

The interesting use is a small check placed before an expensive or consequential action. A routing question does not necessarily need another long generated response. But a fixed choice and a confidence score still need a surrounding policy: what happens when the model is wrong, and which decisions require a person? The weather example makes that boundary visible. We would judge this release by measured workflow errors and latency, rather than assuming that a small local model automatically becomes a reliable guardrail.