Back to AI Research

AI Research

AI vs Human Expert Reasoning: Assessing Agreements... | AI Research

Key Takeaways

  • AI vs Human Expert Reasoning: Assessing Agreements in Building Typology Predictions based on Street View Imagery This research explores the effectiveness of...
  • This research investigates the potential of Vision-Language Models (VLMs) to infer building typologies: Construction, Current Use, and Storeys from Google Street View (GSV) images.
  • Predictions generated by VLMs are compared with inference by human experts (civil engineers and architects) as a source of manually labelled ground-truth data.
  • We evaluate several state-of-the-art VLMs, including GPT-4o, Claude 3.5 Sonnet, and Gemini 2.0 Flash.
  • By applying different scaling strategies and prompting techniques, we found that Chain-of-Thought prompts provide an overall more stable model performance.
Paper AbstractExpand

This research investigates the potential of Vision-Language Models (VLMs) to infer building typologies: Construction, Current Use, and Storeys from Google Street View (GSV) images. Predictions generated by VLMs are compared with inference by human experts (civil engineers and architects) as a source of manually labelled ground-truth data. We evaluate several state-of-the-art VLMs, including GPT-4o, Claude 3.5 Sonnet, and Gemini 2.0 Flash. By applying different scaling strategies and prompting techniques, we found that Chain-of-Thought prompts provide an overall more stable model performance. We also investigate the reasoning behind VLMs' building-typology predictions by examining the probabilities of keywords appearing in AI explanations. This enabled us to analyse patterns in these reasonings and identify key themes driving both agreements and disagreements between VLM and expert labels. We find that AI tends to focus on visual indicators, whereas human experts place greater emphasis on broader contextual cues and domain knowledge, in addition to visual cues. Overall, VLM can approximate experts' capability in building-typology classification at scale, with an average accuracy of approximately 70%. The study demonstrates the VLM's potential for AI automation in tasks that require pattern recognition and object identification in an urban context. AI have the potential to serve as complementary and collaborative tools for urban analysis, leveraging their strengths in understanding visual patterns. This study contributes to the exploration of the efficiency and scalability of AI visual prediction and provides insights into the reasoning processes that could support automation processes in urban analysis and prediction.

AI vs Human Expert Reasoning: Assessing Agreements in Building Typology Predictions based on Street View Imagery

This research explores the effectiveness of Vision-Language Models (VLMs) in identifying key building characteristics—specifically construction type, current use, and the number of storeys—using Google Street View imagery. By comparing AI predictions against ground-truth data provided by civil engineers and architects, the study evaluates how well modern AI can replicate expert-level urban analysis and identifies the differences in how humans and machines "reason" about the built environment.

Evaluating Modern Vision-Language Models

The researchers tested several state-of-the-art models, including GPT-4o, Claude 3.5 Sonnet, and Gemini 2.0 Flash. To optimize performance, the team experimented with various scaling strategies and prompting techniques. They discovered that using "Chain-of-Thought" prompting—a method that encourages the AI to break down its reasoning process step-by-step—resulted in more stable and reliable model performance across the board.

How AI and Humans Differ in Reasoning

A core component of the study involved analyzing the "why" behind the AI’s predictions by examining the keywords and themes present in its explanations. The findings reveal a distinct divide in methodology:

  • AI focus: Models primarily rely on direct visual indicators found within the imagery.

  • Human focus: Experts integrate these visual cues with broader contextual knowledge and domain-specific expertise.
    These differences highlight that while AI is highly capable of pattern recognition, it lacks the deep, contextual understanding that human professionals bring to urban analysis.

Performance and Future Potential

The study found that VLMs can achieve an average accuracy of approximately 70% when classifying building typologies. This suggests that AI is already capable of approximating expert-level performance at a scale that would be difficult for humans to achieve manually. The authors conclude that VLMs are best positioned as collaborative tools, serving as a scalable, automated assistant that can handle routine visual identification tasks, thereby allowing human experts to focus on more complex, context-heavy urban planning and analysis.

Comments (0)

No comments yet

Be the first to share your thoughts!