Researchers have developed the Institutional Newspapers Pipeline, a modular system designed to extract high-quality, structured data from historical newspaper scans. This pipeline, created in collaboration with the Boston Public Library, processes scans through segmentation, OCR, and various text analysis steps to classify types, detect reading order, recognize named entities, and generate embeddings. The system was applied to over 1.4 million public domain newspaper scans published between 1795 and 1930, yielding 16.3 billion tokens and resulting in an open dataset. AI
IMPACT Enables large-scale analysis of historical text data, potentially uncovering new insights from digitized archives.
RANK_REASON The item is a research paper describing a new pipeline for data extraction. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- Boston Public Library
- Connected Papers
- DagsHub
- Gotit.pub
- Hugging Face
- Institutional Newspapers Pipeline
- Litmaps
- Matteo Cargnelutti
- optical character recognition
- ScienceCast
- scite Smart Citations
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →