Google Document AI
PulseAugur coverage of Google Document AI — every cluster mentioning Google Document AI across labs, papers, and developer communities, ranked by signal.
-
New Sinhala OCR Dataset and Model Achieve State-of-the-Art Performance
Researchers have developed a new dataset, sinhala-ocr-lk-acts-1010, to improve Optical Character Recognition (OCR) for the Sinhala language, which is spoken by approximately 16 million people in Sri Lanka. This dataset …
-
New Sinhala OCR Dataset and LightOnOCR-2-1B Achieve State-of-the-Art Performance
Researchers have developed a new dataset, sinhala-ocr-lk-acts-1010, to improve Optical Character Recognition (OCR) for Sinhala, a language spoken by approximately 16 million people. This dataset comprises 1,010 page-lev…
-
In the Arena: How LMSys changed LLM Benchmarking Forever
The AraGen benchmark, developed by Hugging Face, aims to improve LLM evaluation by addressing limitations of static benchmarks. It introduces a crowdsourced approach similar to LMSys's Chatbot Arena, allowing for more d…