Researchers have developed news-crawler-LM, a compact language model designed for efficient and accurate extraction of structured content from news articles. This model, fine-tuned using the Fundus news-crawling library, converts raw HTML into plaintext and structured JSON, including key fields like headline, author, and publication date. Experiments show that news-crawler-LM surpasses existing baselines in HTML-to-Markdown and HTML-to-JSON extraction tasks, demonstrating significant improvements in BLEU and METEOR scores. The project also makes its models and artifacts available to the research community. AI
IMPACT This model could streamline content extraction for news aggregators and researchers, reducing manual effort and improving data quality.
RANK_REASON The cluster contains a research paper detailing a new model release and its performance evaluation. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →