The FineBooks project, a collaboration between Hugging Face and EleutherAI, has evaluated 14 open-source OCR models to improve the quality of historical text data for AI training. The leading model, dots.mocr, achieved 97.6% character accuracy at a cost of less than $2 per thousand pages. While this accuracy is sufficient for AI training, the project notes it is not yet precise enough for academic transcription purposes. AI
IMPACT Improves the quality and accessibility of historical text data for training large language models.
RANK_REASON Research project evaluating OCR models for AI training data. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →