A new benchmark called DataSpace has been released, designed to evaluate data agents on their ability to produce verifiable tabular results from diverse workspaces. The benchmark comprises 410 cross-language tasks involving over 7,439 artifacts totaling 15.01 GB across various formats like CSV, JSON, SQLite, Markdown, PDF, and video. Initial tests with frontier multimodal models and agent harnesses show that the choice of harness significantly impacts accuracy, with the best performance reaching 66.34% and swapping harnesses yielding a 15.36-point difference. AI
IMPACT Highlights the critical role of agent harness selection in achieving accurate data retrieval from complex, multimodal sources.
RANK_REASON The cluster describes a new benchmark and research paper, fitting the research bucket. [lever_c_demoted from research: ic=1 ai=1.0]
Read on X — Omar Sanseviero (HF research) →
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →