Researchers have developed an open-source pipeline called Institutional Books - Visual Elements, designed to extract, classify, deduplicate, and caption visual components from digitized historical book collections. This pipeline, along with an initial dataset of 22.6 million visual elements, has been released to facilitate new applications for digitized library materials. The project aims to make visual elements like illustrations and photographs more accessible for computational use, including AI model training and digital humanities research. AI
IMPACT Enables new use cases for digitized library collections through computational access, including AI model training.
RANK_REASON The item describes an open-source pipeline and dataset release for processing visual elements in digitized books, which falls under research and development in computational access to cultural heritage. [lever_c_demoted from research: ic=1 ai=0.7]
- arXiv
- Hugging Face
- Institutional Books: Harvard Library
- Institutional Books - Visual Elements
- Matteo Cargnelutti
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →