Researchers have introduced SoM-1K, a new multimodal benchmark dataset designed to evaluate the capabilities of foundation models on complex engineering problems in the strength of materials. The dataset comprises 1,065 problems, each with textual descriptions and schematic diagrams. A novel prompting strategy, Descriptions of Images (DoI), which uses expert-generated text descriptions of diagrams, was employed to test eight leading large language and vision-language models. The results indicate that current models struggle significantly with these tasks, with the best model achieving only 56.6% accuracy, highlighting a need for improved multimodal reasoning in AI for scientific and engineering applications. AI
IMPACT Highlights a critical need for developing more robust multimodal reasoning capabilities in foundation models for scientific and engineering contexts.
RANK_REASON The cluster describes a new academic paper introducing a benchmark dataset and evaluation of existing models. [lever_c_demoted from research: ic=1 ai=1.0]
- arXiv
- Descriptions of Images
- Hugging Face
- large language models
- Qixin Wang
- SoM-1K
- Vision--Language Models
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →