Social Sciences DataLab
PulseAugur coverage of Social Sciences DataLab — every cluster mentioning Social Sciences DataLab across labs, papers, and developer communities, ranked by signal.
1 day(s) with sentiment data
-
Chandra OCR Model Dominates PDF Parsing Benchmark
A recent comparison of PDF parsing capabilities revealed that Chandra, an OCR model from Datalab, outperformed all other tested parsers, successfully handling merged-cell HTML tables, LaTeX, and even difficult cursive t…
-
RAG parsing of scientific papers struggles with tables and equations
Parsing scientific papers for retrieval-augmented generation (RAG) systems remains challenging due to complex layouts, equations, and tables that often result in extraction errors. These errors, such as incorrect number…
-
Marker 2 document converter achieves 5x throughput, beats competitors on benchmark
Datalab has released Marker 2, a significantly rewritten open-source document conversion pipeline. The new version boasts a 5x increase in throughput compared to MinerU, achieving 2.9 pages per second on a single Nvidia…
-
Datalab's Lift 9B model leads in schema-first PDF extraction
Datalab's Lift is a new 9-billion parameter vision-language model designed for schema-first document extraction. Unlike traditional methods that first parse documents into intermediate formats before extracting fields, …
-
Open-source models tackle PDF-to-JSON conversion for enterprise AI
New open-source models are emerging to convert unstructured data within PDFs into usable JSON formats, addressing a critical need for enterprise AI applications. These models fall into two main categories: schema-driven…
-
Open-source OCR models and benchmarks consolidated on Papers with Code
A new resource has been created to track open-source optical character recognition (OCR) models, consolidating information on top-performing models, benchmarks, and links to their papers and code. This initiative highli…
-
Datalab releases lift, a 9B open-weights vision model for structured PDF extraction
Datalab has launched lift, a 9B parameter open-weights vision model designed for structured data extraction from PDFs and images. The model takes a JSON schema as input and generates a JSON object conforming to that sch…