Researchers have developed a method to improve Optical Character Recognition (OCR) for low-resource historical languages, specifically focusing on Manchu. By combining synthetic and real historical word images, they achieved a word accuracy of up to 96.28% on Qing dynasty manuscripts. The study found that joint and sequential training methods yielded similar results, and a compact Convolutional Recurrent Neural Network (CRNN) also achieved high performance when real images were incorporated. Complementary error analysis and dictionary-based voting further boosted accuracy to 98.27% without additional training. AI
IMPACT Improves accessibility and searchability of historical archives for endangered languages.
RANK_REASON Academic paper detailing a new method for OCR on a historical language. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →