AI vs Human Expert Reasoning: Assessing Agreements in Building Typology Predictions based on Street View Imagery
This research explores the effectiveness of Vision-Language Models (VLMs) in identifying key building characteristics—specifically construction type, current use, and the number of storeys—using Google Street View imagery. By comparing AI predictions against ground-truth data provided by civil engineers and architects, the study evaluates how well modern AI can replicate expert-level urban analysis and identifies the differences in how humans and machines "reason" about the built environment.
Evaluating Modern Vision-Language Models
The researchers tested several state-of-the-art models, including GPT-4o, Claude 3.5 Sonnet, and Gemini 2.0 Flash. To optimize performance, the team experimented with various scaling strategies and prompting techniques. They discovered that using "Chain-of-Thought" prompting—a method that encourages the AI to break down its reasoning process step-by-step—resulted in more stable and reliable model performance across the board.
How AI and Humans Differ in Reasoning
A core component of the study involved analyzing the "why" behind the AI’s predictions by examining the keywords and themes present in its explanations. The findings reveal a distinct divide in methodology:
AI focus: Models primarily rely on direct visual indicators found within the imagery.
Human focus: Experts integrate these visual cues with broader contextual knowledge and domain-specific expertise.
These differences highlight that while AI is highly capable of pattern recognition, it lacks the deep, contextual understanding that human professionals bring to urban analysis.
Performance and Future Potential
The study found that VLMs can achieve an average accuracy of approximately 70% when classifying building typologies. This suggests that AI is already capable of approximating expert-level performance at a scale that would be difficult for humans to achieve manually. The authors conclude that VLMs are best positioned as collaborative tools, serving as a scalable, automated assistant that can handle routine visual identification tasks, thereby allowing human experts to focus on more complex, context-heavy urban planning and analysis.
Comments (0)
to join the discussion
No comments yet
Be the first to share your thoughts!