Researchers have developed the Institutional Newspapers Pipeline, a modular system designed to extract high-quality, structured data from historical newspaper scans. This pipeline, created in collaboration with the Boston Public Library, segments scans, performs optical character recognition (OCR), and then analyzes text for various features like named entities and subject classification. The resulting open dataset includes over 16 billion tokens from nearly 1.5 million newspaper scans published between 1795 and 1930, making historical public life records more computationally accessible. AI
IMPACT Enables new forms of historical research and data analysis by making vast archives of text computationally accessible.
RANK_REASON The item describes a research paper detailing a new method and dataset for processing historical documents. [lever_c_demoted from research: ic=1 ai=0.7]
Read on Hugging Face Daily Papers →
- Boston Public Library
- Harvard Law Library
- Institutional Data Initiative
- Institutional Newspapers Pipeline
- o200k_base
- optical character recognition
- Tesseract
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →