Researchers have introduced Dashboard2Code, a new task designed to evaluate multimodal large language models' ability to reconstruct interactive dashboards. This task goes beyond static chart generation by requiring models to interact with dashboards, integrate feedback, and produce code that replicates the target dashboard. To facilitate this, they developed DashboardMimic, a benchmark dataset using Plotly+Dash with 180 dashboard-code pairs across varying difficulty levels and interaction patterns. Experiments show that current models, including strong open- and closed-source options, struggle with complex dashboards, highlighting a significant performance gap. AI
IMPACT This research highlights limitations in current multimodal models for complex interactive data visualization, suggesting areas for future development.
RANK_REASON The cluster describes a new research paper introducing a novel task and benchmark for evaluating multimodal models.
AI-generated summary · Google Gemini · from 3 sources. How we write summaries →