A new benchmark called GTPred has been introduced to evaluate the geo-temporal prediction capabilities of multi-modal large language models (MLLMs). This benchmark includes 370 images from over 120 years and assesses MLLMs' ability to infer both location and time of capture. Experiments with 15 different MLLMs revealed that while current models excel at visual perception, they struggle with world knowledge and geo-temporal reasoning. The study also indicated that incorporating temporal information significantly improves location inference accuracy. AI
IMPACT Highlights limitations in current MLLMs' world knowledge and geo-temporal reasoning, suggesting areas for future development.
RANK_REASON The cluster describes a new academic benchmark and evaluation of existing models. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →