A new study published on arXiv challenges the assumption that vision models accurately represent urban scenes as humans do. Researchers found that while models like DINOv2 ViT-B can predict human appraisal ratings of street scenes with high accuracy (up to r = 0.87), their internal representations do not align with neural data from electroencephalography (EEG) recordings of human perception. The best-performing model only explained 29.6% of the explainable neural geometry, and this alignment did not improve appraisal prediction. This suggests that high predictive accuracy in rating tasks may not be a reliable indicator of how well these models understand visual scenes. AI
IMPACT Challenges the reliability of current benchmarks for evaluating vision models' understanding of visual scenes.
RANK_REASON The cluster contains a research paper published on arXiv detailing findings about AI model representations. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- Berlin
- CatalyzeX
- DagsHub
- DINOv2 ViT-B
- electroencephalography
- Gotit.pub
- Hugging Face
- ScienceCast
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →