VIALS is a new benchmark designed to evaluate how well vision-language models (VLMs) interpret the specialized visual artifacts used in professional life sciences research. While current AI models can describe natural images, researchers at Handshake AI found that these models struggle to accurately process scientific data—such as gel blots, plasmid maps, and flow cytometry plots—that are essential for decision-making in the biotech industry.
The VIALS Benchmark
The benchmark consists of 161 visual question-answering tasks that reflect real-world experimental workflows rather than polished textbook figures. These tasks were created and vetted by 31 PhD-level scientists with professional experience in fields like immunology, structural biology, and medicinal chemistry. Each task requires a model to extract specific evidence from an image and apply scientific reasoning to provide a correct answer. The dataset is designed to be challenging, featuring "messy" experimental data that scientists encounter daily in their work.
How Models Perform
When tested against the benchmark, top-performing models like GPT-5.6 Sol and Gemini 3.7 Flash achieved only 26.5% accuracy. The researchers observed that these models frequently fail to perform basic tasks, such as counting objects, reading measurements, or understanding the spatial relationships within a scientific plot. Even when a model correctly identifies a feature, it may still apply the wrong scientific convention or rely on prior expectations rather than the actual evidence present in the image.
Understanding Failure Modes
The researchers categorized model errors to identify why these systems struggle. The most common failure, accounting for up to 41.8% of errors, is visual quantification—the inability to accurately count cells, colonies, or bands. Other frequent issues include selecting the wrong feature, overlooking relevant evidence, and misinterpreting the structural conventions of scientific representations. Fabricated findings, where a model invents data not present in the image, were found to be rare.
Why This Matters
The authors argue that for AI to be useful in professional life sciences, it must be able to reliably interpret the visual artifacts that drive research progress. Because current models fail to reach high accuracy levels on these tasks, they cannot yet be trusted for high-stakes scientific decision-making. The VIALS benchmark provides a standardized way to measure progress in this area, ensuring that future AI development is grounded in the specific, high-value needs of the biotechnology and pharmaceutical industries.
Comments (0)
to join the discussion
No comments yet
Be the first to share your thoughts!