A new research paper explores the capabilities of Multimodal Large Language Models (MLLMs) in assessing perceived urban safety from street-view imagery. While these models demonstrate a zero-shot capability to predict safety with reasonable accuracy across various cities, they exhibit a tendency to favor 'Safe' classifications and underpredict unsafety. Furthermore, the study reveals that MLLMs encode non-neutral demographic priors, showing significant shifts in safety perception when prompted with specific gender, age, or racial/ethnic personas. AI
IMPACT Reveals that MLLMs can be used to assess urban safety but carry demographic biases, impacting their neutrality in planning applications.
RANK_REASON Research paper published on arXiv detailing findings about MLLMs. [lever_c_demoted from research: ic=1 ai=1.0]
- age
- Ciro Beneduce
- Female
- gender
- Native American
- Male
- Multimodal Large Language Models
- Place Pulse 2.0
- race or ethnicity
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →