Researchers have developed a new benchmark called URBANCONTRASTIVEQA to evaluate how well tool-augmented large language models can understand relative urban activity. The benchmark presents pairs of urban situations from public mobility data in New York City, Chicago, and Seattle, asking models to determine which scenario is more abnormal relative to its local historical baseline, rather than just picking the larger raw count. Results show that models often struggle with this comparative task when only given raw counts, but accuracy improves when baseline scores and ordinal labels are provided, though gains vary by model. AI
IMPACT This benchmark could lead to more nuanced LLM capabilities in analyzing real-world data, improving decision-support tools.
RANK_REASON The cluster contains a research paper detailing a new benchmark for evaluating LLMs. [lever_c_demoted from research: ic=1 ai=1.0]
Read on arXiv cs.IR (Information Retrieval) →
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →