Researchers have developed HTML-LM, a compact foundation model with 154 million parameters designed to efficiently process Czech HTML documents. This model, built on a ModernBERT architecture, utilizes HTML-aware training and multiple objectives, including masked language modeling and contrastive distillation from larger models. Deployed in production, HTML-LM processes thousands of web documents per second and achieves state-of-the-art performance for classification and regression tasks within the Czech internet domain, outperforming larger encoders and smaller LLMs. AI
IMPACT This model's efficiency and performance in processing structured web data could inform future developments in web crawling, information extraction, and domain-specific language models.
RANK_REASON The cluster describes the release of a new academic paper detailing a novel model architecture and its performance on a specific domain. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →