Back to AI Research

AI Research

Mingbird tests whether better scaffolding helps small local models finish tasks

Key Takeaways

  • A Windows-and-Ollama agent harness targets false finishes, repeated tool calls and oversized prompts, with reported gains in a controlled single-machine experiment.
  • A local model can choose a sensible next action yet still fail to finish a task.
  • It may repeat a tool call, lose the original acceptance criteria or declare success before producing the required file.
  • Mingbird's authors treat those failures as problems in the software surrounding the model and test whether targeted changes improve completion.
  • Their [Mingbird research paper](https://arxiv.org/abs/2610.02001) describes a Windows-and-Ollama agent harness intended for small open-weight models.

A local model can choose a sensible next action yet still fail to finish a task. It may repeat a tool call, lose the original acceptance criteria or declare success before producing the required file. Mingbird's authors treat those failures as problems in the software surrounding the model and test whether targeted changes improve completion.
Their Mingbird research paper describes a Windows-and-Ollama agent harness intended for small open-weight models. The contribution is systems engineering rather than a new learning algorithm. Its reported results are promising within the evaluated setup, but the authors acknowledge a self-built benchmark, one machine and single-trial scoring.

A smaller prompt and a stricter finish

Mingbird routes tasks to relevant tool definitions instead of loading the entire tool surface into every request. Its documented v1.6.0 release fixes the estimated static prefill at 797 tokens, with a regression check requiring new features to add zero net static prompt growth. The benchmark used v1.5.0, whose corresponding estimate was 774 tokens. These are harness estimates, separate from backend-reported prompt measurements.
Routing has a cost: a mistaken category can hide a required tool. The authors therefore allow re-routing. They also distinguish casual chat from artifact-producing tasks so a simple question does not invite a small model to demonstrate tools in an open-ended loop.
Before accepting a completion claim, the harness rereads the original task and checks delivery against it. Real check failures return to the conversation. That adds a model turn and cannot guarantee a correct self-check, but it gives the agent another opportunity with the actual requirements in context.
Repeated tool calls receive their own guard. Mingbird normalizes the tool name and arguments into a signature, then uses thresholds for identical calls and consecutive calls without output. The authors describe graduated corrective feedback rather than terminating a run at the first repetition. File backups and resumable checkpoints address other failures outside model reasoning.

Holding the rest of the experiment constant

The authors built LRAB to vary the harness while holding machine, models, backend, budgets and scoring fixed. The experiment covers four harnesses, four open models spanning 2B to 35B parameters and 18 tasks, producing 288 scored cells. Artifact-based scoring checks the output rather than accepting a model's claim that it completed the task.
Mingbird's reported overall score is 0.886, compared with 0.631 for goose, 0.479 for opencode and 0.405 for agent-mini. These figures describe the published LRAB protocol, not general product rankings or performance on an arbitrary laptop.
The paper also reports a three-arm evaluation on tau-squared-bench. Across 278 scored retail, airline and telecom tasks, Mingbird scores 0.856, against 0.791 for the benchmark's native agent and 0.737 for opencode. The goose arm was deferred for throughput. Although the harness is local-first, this evaluation used a cloud-hosted user simulator and judge.

Evidence for individual mechanisms remains narrower

The paper warns against treating its single-trial ablations as precise causal estimates. Repeating the same arm on the same night moved its mean by as much as 0.069, comparable to nominal single-trial mechanism differences. A batch-matched comparison across three replications reports a paired gain of 0.10 for the full completion guards over text rereading alone.
For developers building local agents, Mingbird offers concrete failure-handling patterns to examine. The study supports testing prompt budgets, completion checks and loop recovery alongside model choice. It does not establish that scaffolding removes model capability limits or that the measured gains will transfer unchanged to other workloads.

Comments