Back to AI Research

AI Research

VIALS: A Benchmark for Visual Interpretation of Art... | AI Research

Key Takeaways

  • VIALS is a new benchmark designed to evaluate how well vision-language models (VLMs) interpret the specialized visual artifacts used in professional life sci...
  • In professional life sciences workflows, scientists routinely interpret visual artifacts (gel blots, microscopy images, plasmid maps, flow cytometry plots, molecular structures, ...) to inform research decisions.
  • In contrast, scientists with relevant domain expertise find these visual interpretation tasks straightforward.
  • AI that cannot similarly interpret such images will have limited utility in professional life sciences workflows, where such artifacts are central to how scientists reason, communicate, and make decisions.
  • VIALS is a new benchmark designed to evaluate how well vision-language models (VLMs) interpret the specialized visual artifacts used in professional life sciences research.
Paper AbstractExpand

In professional life sciences workflows, scientists routinely interpret visual artifacts (gel blots, microscopy images, plasmid maps, flow cytometry plots, molecular structures, ...) to inform research decisions. We introduce VIALS, a visual question-answering benchmark with 161 such interpretation tasks, spanning the types of artifacts examined throughout experimental workflows in the biotech industry (rather than polished figures from publications and textbooks). While frontier vision-language models can now fluently describe natural images, we find that they are unable to accurately interpret these scientific images, reflecting limitations in domain knowledge and domain-specific visual reasoning capabilities. In contrast, scientists with relevant domain expertise find these visual interpretation tasks straightforward. AI that cannot similarly interpret such images will have limited utility in professional life sciences workflows, where such artifacts are central to how scientists reason, communicate, and make decisions.

VIALS is a new benchmark designed to evaluate how well vision-language models (VLMs) interpret the specialized visual artifacts used in professional life sciences research. While current AI models can describe natural images, researchers at Handshake AI found that these models struggle to accurately process scientific data—such as gel blots, plasmid maps, and flow cytometry plots—that are essential for decision-making in the biotech industry.

The VIALS Benchmark

The benchmark consists of 161 visual question-answering tasks that reflect real-world experimental workflows rather than polished textbook figures. These tasks were created and vetted by 31 PhD-level scientists with professional experience in fields like immunology, structural biology, and medicinal chemistry. Each task requires a model to extract specific evidence from an image and apply scientific reasoning to provide a correct answer. The dataset is designed to be challenging, featuring "messy" experimental data that scientists encounter daily in their work.

How Models Perform

When tested against the benchmark, top-performing models like GPT-5.6 Sol and Gemini 3.7 Flash achieved only 26.5% accuracy. The researchers observed that these models frequently fail to perform basic tasks, such as counting objects, reading measurements, or understanding the spatial relationships within a scientific plot. Even when a model correctly identifies a feature, it may still apply the wrong scientific convention or rely on prior expectations rather than the actual evidence present in the image.

Understanding Failure Modes

The researchers categorized model errors to identify why these systems struggle. The most common failure, accounting for up to 41.8% of errors, is visual quantification—the inability to accurately count cells, colonies, or bands. Other frequent issues include selecting the wrong feature, overlooking relevant evidence, and misinterpreting the structural conventions of scientific representations. Fabricated findings, where a model invents data not present in the image, were found to be rare.

Why This Matters

The authors argue that for AI to be useful in professional life sciences, it must be able to reliably interpret the visual artifacts that drive research progress. Because current models fail to reach high accuracy levels on these tasks, they cannot yet be trusted for high-stakes scientific decision-making. The VIALS benchmark provides a standardized way to measure progress in this area, ensuring that future AI development is grounded in the specific, high-value needs of the biotechnology and pharmaceutical industries.

Comments (0)

No comments yet

Be the first to share your thoughts!