This paper investigates whether AI agents can perform open-ended research, a capability often cited as a prerequisite for explosive progress in AI development. The authors argue that current evaluation methods—such as testing agents on narrow, verifiable tasks or using inconsistent blind peer reviews—are insufficient for measuring true research autonomy.
The Shadow Evaluation Method
To address these limitations, the researchers introduced "shadow evaluations." In this framework, an AI agent is tasked with answering the core, open-ended research question from a high-quality, unpublished paper. The original authors of those papers then grade the agent's output. This approach aims to provide a more rigorous assessment of an agent's ability to navigate the complexities of the research lifecycle compared to existing benchmarks.
Experimental Results
The team conducted shadow evaluations on two unpublished NeurIPS 2026 submissions. They provided frontier AI agents with six days and thousands of dollars in compute resources to complete the research. While the agents successfully handled the engineering components of the projects without human intervention, they failed to make substantial progress toward answering the primary research questions. Consequently, the original authors rejected both agent-generated papers. A robustness check using a second model and scaffold confirmed these findings.
Recurring Failure Modes
The researchers identified five specific areas where the agents struggled:
Judgment: Difficulty assessing the standard required for publishable research.
Creativity: Inability to generate effective responses to research design flaws.
Backtracking: Failure to recover from dead ends during the research process.
Resource Awareness: Poor management of compute and time resources.
Instruction Drift: A tendency to lose focus on the original research objectives.
Implications for AI Research
The study concludes that while current AI agents are capable of performing the engineering tasks associated with research, they lack the critical skills necessary to manage the broader research lifecycle. These findings suggest that significant gaps remain in the ability of AI to conduct autonomous, open-ended scientific inquiry. The authors have released the expert reviews, survey responses, agent repositories, and logs to support further investigation into these limitations.
Comments (0)
to join the discussion
No comments yet
Be the first to share your thoughts!