A new benchmark called DataSpace has been introduced to evaluate data agents' ability to perform verifiable analytics over complex, heterogeneous workspaces. The benchmark includes 410 tasks and over 7,000 artifacts totaling 15 GB across various formats like CSV, JSON, SQLite, Markdown, PDF, and video. DataSpace was also the official evaluation platform for the KDD Cup 2026 competition, challenging participants to develop agents capable of discovering evidence, integrating information across formats, and producing verifiable tabular results. Current frontier multimodal models achieve a maximum accuracy of 66.34% on DataSpace, highlighting significant challenges in multimodal evidence integration and cross-source joins. AI
IMPACT Highlights current limitations in AI agent reliability for complex data analysis, indicating a need for improved multimodal integration and cross-source joining capabilities.
RANK_REASON The cluster describes a new benchmark and dataset for evaluating AI agents, presented in an academic paper.
Read on Hugging Face Daily Papers →
- alphaXiv
- CatalyzeX
- CSV
- DagsHub
- DataSpace-Builder
- Gotit.pub
- Hugging Face
- JSON
- KDD Cup 2026
- Markdown
- ScienceCast
- SQLite
- video recording
- Data Agents for Complex Data Analysis
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →