VIALS is a new benchmark designed to evaluate how well vision-language models (VLMs) interpret the specialized visual artifacts used in professional life sciences research. While current AI models can describe natural images, researchers at Handshake AI found that these models struggle to accurately process scientific data—such as gel blots, plasmid maps, and flow cytometry plots—that are essential for decision-making in the biotech industry. Methods and results are detailed in the full paper on arxiv.org.
The VIALS Benchmark
The benchmark consists of 161 visual question-answering tasks that reflect real-world experimental workflows rather than polished textbook figures. These tasks were created and vetted by 31 PhD-level scientists with professional experience in fields like immunology, structural biology, and medicinal chemistry. Each task requires a model to extract specific evidence from an image and apply scientific reasoning to provide a correct answer. The dataset is designed to be challenging, featuring "messy" experimental data that scientists encounter daily in their work. The same AI Evaluation question is explored in SciMIF, which adds a research perspective.
How Models Perform When tested against the benchmark, top-performing models like GPT-5.6 Sol and Gemini 3.7 Flash achieved only 26.5% accuracy.
The researchers observed that these models frequently fail to perform basic tasks, such as counting objects, reading measurements, or understanding the spatial relationships within a scientific plot. Even when a model correctly identifies a feature, it may still apply the wrong scientific convention or rely on prior expectations rather than the actual evidence present in the image. The same AI Evaluation question is explored in MNIST-PRO, which adds a research perspective.
Understanding Failure Modes
The researchers categorized model errors to identify why these systems struggle. The most common failure, accounting for up to 41.8% of errors, is visual quantification—the inability to accurately count cells, colonies, or bands. Other frequent issues include selecting the wrong feature, overlooking relevant evidence, and misinterpreting the structural conventions of scientific representations. Fabricated findings, where a model invents data not present in the image, were found to be rare. For a practical look at visual, Flux3.studio is a useful comparison.
Comments