Researchers have developed a novel pipeline to extract and geocode commercial advertisements from digitized 20th-century Armenian newspapers in France. This system utilizes vision-language models (VLMs) to overcome the challenges of processing the under-resourced Western Armenian language and the poor quality of scanned historical documents. The project introduces a new corpus of Western Armenian press with advertisement annotations and a reproducible workflow applicable to similar historical language data. AI
IMPACT Demonstrates VLM effectiveness for under-resourced historical languages, potentially enabling new research avenues in digital humanities.
RANK_REASON The item is an academic paper detailing a new methodology for processing historical documents using AI. [lever_c_demoted from research: ic=1 ai=1.0]
- Armenian Paris
- Chahan Vidal-Gorène
- Convolutional Recurrent Neural Network
- France
- Hugging Face
- Label Studio
- Western Armenian
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →