PulseAugur
EN
LIVE 18:15:58

AI models approximate human experts in building typology predictions from street view

A new research paper evaluates the capabilities of Vision-Language Models (VLMs) in predicting building typologies from Google Street View imagery. The study compares the performance of models like GPT-4o, Claude 3.5 Sonnet, and Gemini 2.0 Flash against human experts, finding that VLMs achieve approximately 70% accuracy. While VLMs focus on visual cues, human experts incorporate broader contextual knowledge, suggesting VLMs can serve as scalable tools for urban analysis. AI

IMPACT VLMs can approximate expert capabilities in urban analysis tasks, offering scalable automation for pattern recognition and object identification.

RANK_REASON The cluster contains an academic paper detailing research findings on AI model capabilities.

Read on arXiv cs.AI →

AI-generated summary · Google Gemini · from 2 sources. How we write summaries →

AI models approximate human experts in building typology predictions from street view

COVERAGE [2]

  1. arXiv cs.AI TIER_1 English(EN) · Zahratu Shabrina, Muhammad Asa, Jin Rui, Lu Yin, Stephen Law ·

    AI vs Human Expert Reasoning: Assessing Agreements in Building Typology Predictions based on Street View Imagery

    arXiv:2607.14756v1 Announce Type: new Abstract: This research investigates the potential of Vision-Language Models (VLMs) to infer building typologies: Construction, Current Use, and Storeys from Google Street View (GSV) images. Predictions generated by VLMs are compared with inf…

  2. arXiv cs.AI TIER_1 English(EN) · Stephen Law ·

    AI vs Human Expert Reasoning: Assessing Agreements in Building Typology Predictions based on Street View Imagery

    This research investigates the potential of Vision-Language Models (VLMs) to infer building typologies: Construction, Current Use, and Storeys from Google Street View (GSV) images. Predictions generated by VLMs are compared with inference by human experts (civil engineers and arc…