Back to AI Research

AI Research

OSWorld-Science tests whether computer agents produce valid scientific outputs

Key Takeaways

  • A 146-task benchmark checks actual scientific files and application states, revealing frequent missing deliverables despite partial-credit performance gains.
  • An agent can navigate scientific software and still finish without a usable result.
  • The benchmark contains 146 scientific software tasks and evaluates twelve visual language models.
  • Workflows include molecular drawing, pathology image analysis, statistical computing and physical simulation.
  • Experts propose tasks, with additional tasks developed through iterative human-AI collaboration.

An agent can navigate scientific software and still finish without a usable result. OSWorld-Science, introduced by Dingyuan Dai and colleagues, evaluates the files and application states that agents produce rather than accepting a confident description of completed work.
The benchmark contains 146 scientific software tasks and evaluates twelve visual language models. Workflows include molecular drawing, pathology image analysis, statistical computing and physical simulation. Experts propose tasks, with additional tasks developed through iterative human-AI collaboration. Selection considers scientific value and difficulty.

Evaluators inspect the scientific output

Task-specific evaluators examine molecular structures, segmentation masks, plots or numerical results, depending on the goal. Scores range from zero to one, with partial credit for incomplete outcomes. A high mean score therefore differs from the proportion of tasks completed in full.
Agents operate an Ubuntu environment through graphical actions and, where permitted, command-line interaction. All tasks require screenshot-based observation of the scientific application state. The recorded configurations permit command-line access for 122 tasks, while 24 are GUI-only. That makes interface knowledge and correct manipulation of scientific objects part of the test, rather than reducing each task to a question-answering exercise.
The harness records model requests, responses and executed actions, including return codes and errors. Final grading uses collected files and application state. An agent saying it is done cannot substitute for the required artifact.

Partial credit and missing deliverables tell different stories

The authors report a highest overall mean task-specific score of 73.7% under the evaluated settings. That metric includes partial credit and assigns zero to runs that exhaust their budget without a valid outcome. It should not be read as a 73.7% full-completion rate.
Trajectory analysis provides a separate count: 402 runs, or 26.3%, received full credit. Among 1,100 other runs after excluding harness-ended and unresolved cases, 667 ended without the graded artifact. Of those, 613 exhausted their budget. Missing outputs account for more failures in that subset than incorrect or invalid artifacts.
The benchmark is uneven across domains. Chemistry contributes 43 tasks, while linguistics contributes two. Aggregate scores can hide those differences, and an apparent strength on a tiny domain slice needs a different interpretation from performance across dozens of tasks. Model rankings also change across domains.

More reasoning does not guarantee a better result

The authors examine reasoning effort, screenshot-history length and task language on a 23-task QuPath pathology subset. Increasing reasoning effort raises token use, but task outcomes can improve, regress or remain unchanged. None of the non-default configurations differs significantly from its default after the study's multiple-comparison correction.
Across the wider trajectory analysis, apparent associations between interaction style and success disappear within the same model and domain. More narration or a higher share of command-line actions should therefore not be treated as a proven way to improve a particular agent.
For scientific automation, the useful adoption test is whether the agent produces the correct required output within the available budget. These results motivate checking actual artifacts, reporting full completion alongside partial credit and evaluating the software and domain that a deployment will use. They do not establish clinical readiness or general scientific reliability.

Comments (0)

No comments yet

Be the first to share your thoughts!