Back to AI Research

AI Research

InSituMeasure: Probing Situated Measurement Groundi... | AI Research

Key Takeaways

  • Multimodal Large Language Models (MLLMs) have shown impressive capabilities in general tasks, yet they struggle significantly when tasked with reading indust...
  • For trained operators, gauge reading requires little specialized knowledge, low cognitive effort, and high repeatability.
  • Yet Multimodal Large Language Models (MLLMs) remain unreliable in continuous-valued measurement despite strong results on general multimodal benchmarks.
  • We introduce InSituMeasure to evaluate situated measurement grounding.
  • It contains 2,922 real industrial monitoring scenes across eight functional categories of professional engineering instruments, with dense gauge-attribute annotations and noise tags for failure diagnosis.
Paper AbstractExpand

For trained operators, gauge reading requires little specialized knowledge, low cognitive effort, and high repeatability. Yet Multimodal Large Language Models (MLLMs) remain unreliable in continuous-valued measurement despite strong results on general multimodal benchmarks. Existing benchmarks expose this weakness but isolate measurement from realistic, knowledge-grounded settings, with limited situated context, specialized instruments, real-world noise, and matched diagnostic annotations, reducing realism and constraining root-cause analysis. We introduce InSituMeasure to evaluate situated measurement grounding. It contains 2,922 real industrial monitoring scenes across eight functional categories of professional engineering instruments, with dense gauge-attribute annotations and noise tags for failure diagnosis. We define metrics for numerical accuracy under predefined tolerances and unit consistency, rejection of fake or unanswerable tasks, and alignment between model failures and annotated error factors. Across 24 state-of-the-art MLLMs, the best model reaches only 25.7\% joint value-unit accuracy and 51.8\% confidence-diagnosis F1, revealing a substantial gap between general multimodal competence and reliable situated measurement. Further analysis identifies failures from text-induced shortcuts, overconfident responses, and authentic industrial noise, including mixed disturbances, viewpoint deviation, occlusion, and environmental interference.

Multimodal Large Language Models (MLLMs) have shown impressive capabilities in general tasks, yet they struggle significantly when tasked with reading industrial gauges. While human operators perform these tasks with ease, current AI models lack the reliability needed for real-world engineering environments. The paper InSituMeasure: Probing Situated Measurement Grounding in Industrial Scenes with Multimodal Large Language Models introduces a new benchmark designed to test how well these models can interpret complex, real-world industrial instruments.

Evaluating Real-World Industrial Accuracy

The researchers developed InSituMeasure to bridge the gap between general AI performance and the specialized requirements of industrial monitoring. The dataset consists of 2,922 real-world industrial scenes covering eight categories of professional engineering instruments. Unlike previous benchmarks, this collection includes dense annotations for gauge attributes and specific "noise tags"—such as environmental interference, occlusion, or viewpoint deviations—that help researchers understand why a model might fail to provide an accurate reading.

Defining New Performance Metrics

To move beyond simple accuracy, the authors established a rigorous evaluation framework. They measure success based on numerical accuracy within predefined tolerances and the ability to maintain unit consistency. Furthermore, the benchmark tests the model’s ability to recognize when a task is impossible or "fake," and evaluates how well the model’s internal reasoning aligns with the actual causes of error in the image. This approach allows for a deeper analysis of whether a model is truly "grounding" its measurement in the visual data or simply guessing. The same ai evaluation question is explored in SciMIF, which adds a research perspective.

A Significant Performance Gap

When testing 24 state-of-the-art MLLMs, the results revealed a substantial performance deficit. The top-performing model achieved only 25.7% accuracy in joint value-unit readings and a 51.8% F1 score for confidence-diagnosis. These findings highlight that even the most advanced models currently available are not yet reliable enough for industrial applications, where precision and the ability to account for environmental noise are critical.

Why Models Struggle

The study identifies several core reasons for these failures. Models often rely on "text-induced shortcuts" rather than analyzing the visual evidence, and they frequently exhibit overconfidence in their incorrect answers. Additionally, the models struggle to navigate the complexities of authentic industrial settings, such as mixed visual disturbances, equipment occlusion, and varying camera angles. These factors suggest that current MLLMs require more specialized training to handle the nuances of situated, real-world measurement. The same large language models question is explored in Constrained Entity Selection under Partial Knowledge..., which adds a research perspective. as detailed in the full paper on Arxiv

Comments (0)

No comments yet

Be the first to share your thoughts!