Researchers have introduced MMArt, a new multimodal dataset designed to improve the art interpretation capabilities of vision-language models. Existing datasets offer only single perspectives on artworks, limiting models' ability to perform formal analysis, historical interpretation, or affective characterization. MMArt addresses this by providing 74,234 WikiArt paintings, each annotated with four distinct perspectives—narrative, formal, emotional, and historical—along with a unified caption. Complementary analyses demonstrate that these perspectives encode unique information and are essential for different tasks, such as retrieval and reconstruction, highlighting the value of MMArt's multi-perspective approach. AI
IMPACT This dataset could enable more nuanced AI understanding of art, moving beyond surface-level descriptions to deeper analysis.
RANK_REASON The cluster describes a new academic dataset for AI research. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- CatalyzeX
- DagsHub
- Gotit.pub
- Hugging Face
- Influence Flower
- ScienceCast
- vision-language model
- WikiArt
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →