Researchers have introduced UrbanAgent, a novel framework designed to tackle complex urban tasks by integrating large language models with a suite of tools for code execution and API calls. This system aims to bridge the gap between fragmented digital city services and residents' needs, enabling more seamless cross-system workflows. UrbanAgent demonstrated a 71% task success rate in experiments, outperforming existing baselines by 10 points across various leading LLMs including GPT-5 mini, Gemini 2.5-Flash, DeepSeek-V4 Flash, and Qwen3-235B-A22B. To facilitate evaluation, a new benchmark called Urban-Eval was also developed, which assesses both task outcomes and the quality of execution, including tool coverage and evidence traceability. AI
IMPACT This framework could improve the integration and usability of urban digital services, potentially leading to more efficient city operations and better resident experiences.
RANK_REASON The cluster contains a research paper detailing a new agent framework and benchmark for urban tasks. [lever_c_demoted from research: ic=1 ai=1.0]
- arXiv
- DeepSeek-V4 Flash
- Gemini 2.5-Flash
- GPT-5 mini
- Model Context Protocol
- Qwen3-235B-A22B
- UrbanAgent
- Urban-Eval
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →