Researchers have introduced SCAFFOLD, a novel dataset designed to train vision-language models on understanding complex diagrams found in computer science research papers. This dataset includes images, captions, contextual information, questions, answers, and step-by-step reasoning traces derived from arXiv papers. SCAFFOLD is available in multiple sizes, with the largest version, SCAFFOLD-157K, containing nearly 157,000 question-answer pairs from over 3,000 papers. Initial experiments utilized the smaller SCAFFOLD-12K dataset with the Qwen2.5-VL-3B-Instruct model. AI
IMPACT Enables development of AI models capable of understanding complex visual information in scientific literature.
RANK_REASON The item is a research paper describing a new dataset for AI model training. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- CatalyzeX Code Finder for Papers
- Connected Papers
- CORE Recommender
- DagsHub
- Gotit.pub
- Hugging Face
- Influence Flower
- Litmaps
- Qwen2.5-VL-3B-Instruct
- ScienceCast
- scite Smart Citations
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →