PulseAugur
EN
LIVE 01:07:23
ENTITY Social Sciences DataLab

Social Sciences DataLab

PulseAugur coverage of Social Sciences DataLab — every cluster mentioning Social Sciences DataLab across labs, papers, and developer communities, ranked by signal.

Show in brief
Total · 30d
1
7 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
0
1 over 90d
TIER MIX · 90D
TOPICS
SENTIMENT · 30D

1 day(s) with sentiment data

RECENT · PAGE 1/1 · 7 TOTAL
  1. TOOL · CL_186214 ·

    Chandra OCR Model Dominates PDF Parsing Benchmark

    A recent comparison of PDF parsing capabilities revealed that Chandra, an OCR model from Datalab, outperformed all other tested parsers, successfully handling merged-cell HTML tables, LaTeX, and even difficult cursive t…

  2. TOOL · CL_171032 ·

    RAG parsing of scientific papers struggles with tables and equations

    Parsing scientific papers for retrieval-augmented generation (RAG) systems remains challenging due to complex layouts, equations, and tables that often result in extraction errors. These errors, such as incorrect number…

  3. TOOL · CL_162564 ·

    Marker 2 document converter achieves 5x throughput, beats competitors on benchmark

    Datalab has released Marker 2, a significantly rewritten open-source document conversion pipeline. The new version boasts a 5x increase in throughput compared to MinerU, achieving 2.9 pages per second on a single Nvidia…

  4. TOOL · CL_133747 ·

    Datalab's Lift 9B model leads in schema-first PDF extraction

    Datalab's Lift is a new 9-billion parameter vision-language model designed for schema-first document extraction. Unlike traditional methods that first parse documents into intermediate formats before extracting fields, …

  5. TOOL · CL_125789 ·

    Open-source models tackle PDF-to-JSON conversion for enterprise AI

    New open-source models are emerging to convert unstructured data within PDFs into usable JSON formats, addressing a critical need for enterprise AI applications. These models fall into two main categories: schema-driven…

  6. TOOL · CL_108999 ·

    Open-source OCR models and benchmarks consolidated on Papers with Code

    A new resource has been created to track open-source optical character recognition (OCR) models, consolidating information on top-performing models, benchmarks, and links to their papers and code. This initiative highli…

  7. RESEARCH · CL_107143 ·

    Datalab releases lift, a 9B open-weights vision model for structured PDF extraction

    Datalab has launched lift, a 9B parameter open-weights vision model designed for structured data extraction from PDFs and images. The model takes a JSON schema as input and generates a JSON object conforming to that sch…