A model that answers individual questions well may struggle to build a cheaper system for answering millions of similar questions. Agent in a Bottle introduces BOTTLED to evaluate that conversion.
The authors call it bottling: investing limited resources in a task-specific artifact, then applying the artifact to a complete workload. Across ten models, three tasks and two runs per combination, 48 of 60 runs score below the lower bound of their underlying model's zero-shot 95% confidence interval.
Give the agent the complete workload and a fixed budget
Agents receive an unlabeled workload and choose their approach. They can write a rule-based program, train a smaller model, construct a retrieval index or build a hybrid that routes selected cases back to the original model.
The tasks cover product attribute extraction, query-product relevance classification and AI-text detection. Their datasets contain approximately 4.77 million, 2.62 million and 5.62 million instances, respectively.
Each run uses OpenCode with a single 40GB NVIDIA A100, eight CPU cores and 64GB of RAM. The budget is ten hours and five million weighted tokens. Output tokens count twice as much as uncached input; cached input counts one-tenth as much. The accounting includes the agent's permitted calls to its own evaluated model.
Count artifact construction as part of the cost
The comparison separates upfront construction from applying the solution across the batch. A cheap final classifier can still be expensive to discover, label and train. An agent must allocate its budget between those investments and completing the required predictions.
The zero-shot reference answers one instance at a time. Because evaluating the complete workload would be too expensive, the authors estimate its quality and full-workload cost from a uniform sample of 1,000 instances.
A second baseline spends the same token budget labeling examples with GLM 5.3 Flash, then trains small models using fixed recipes. Thirty-one of the 60 bottling runs underperform the stronger of those two distillation students. Autonomous artifact building therefore needs comparison with a simple planned training procedure, not only with costly repeated frontier calls.
Inspect the quality sacrificed for the savings
On query-product relevance, Opus 5's reported bottled macro-F1 is 0.499, compared with 0.609 for its zero-shot reference. The reported cost is $26.28 versus a projected $17,273 for the full zero-shot workload: about 82% of the reference macro-F1 at roughly 657-times lower reported cost.
Those figures describe a measured quality tradeoff and a projected comparison cost. They are not current pricing advice or evidence that a production dataset will yield the same savings.
The paper also compares this artifact with Jev on the same task. Readers exploring that decision-model approach can inspect the Jev playground walkthrough covering email and content moderation, while keeping the research benchmark separate from a product demonstration.
BOTTLED makes the engineering ability to create reusable solutions measurable. Its results suggest evaluating the delivered batch, construction cost and task-specific error consequences together before substituting an artifact for repeated model calls.
Comments