olmOCR-Bench
PulseAugur coverage of olmOCR-Bench — every cluster mentioning olmOCR-Bench across labs, papers, and developer communities, ranked by signal.
-
Marker 2 document converter achieves 5x throughput, beats competitors on benchmark
Datalab has released Marker 2, a significantly rewritten open-source document conversion pipeline. The new version boasts a 5x increase in throughput compared to MinerU, achieving 2.9 pages per second on a single Nvidia…
-
Youtu-Parsing model accelerates document analysis with novel decoding strategies
Researchers have introduced Youtu-Parsing, a novel document parsing model designed for efficient and high-performance content extraction. The system utilizes a Vision Transformer for feature extraction and a Youtu-LLM-2…
-
General LLMs lead medical decision-making; Infinity-Parser2 tops document parsing benchmarks
A recent survey published on arXiv evaluated 18 large language models (LLMs) for medical applications, finding that general-purpose models outperformed specialized ones in decision-making tasks, while specialist models …
-
Infinity-Parser2 model advances document parsing with synthetic data and multi-task RL
Researchers have introduced Infinity-Parser2, a large multimodal model designed for end-to-end document parsing. The model utilizes a controllable data-synthesis pipeline and multi-task reinforcement learning to overcom…