PulseAugur
实时 21:04:02
English(EN) Where Does the Signal Live? A Web Data Recipe for Medical Encoder Pretraining

新的网络数据配方促进了法国医学编码器预训练

研究人员开发了一种使用网络规模数据预训练医学语言编码器的新方法,解决了较小、手动策划语料库的局限性。他们的方法包括过滤文档以提高医学术语密度,并使用 LLM 将其重写为具有更广泛实体上下文的更密集变体。这种“配方”应用于法国医学 NLP,产生了 FineMed 语料库和 DoctoBERT 编码器系列,在临床命名实体识别任务上表现出最先进的性能。 AI

影响 这项研究可以实现更具可扩展性和多样性的专业领域编码器预训练,有可能提高医学等领域的性能。

排序理由 该集群包含一篇详细介绍语言模型预训练新方法的论文。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.CL 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

新的网络数据配方促进了法国医学编码器预训练

报道来源 [1]

  1. arXiv cs.CL TIER_1 English(EN) · Fajwel Fogel ·

    信号在哪里?用于医学编码器预训练的网页数据配方

    Web data curation has been widely studied for decoder Large Language Model (LLM) pretraining. Encoders for dense-terminology domains such as medicine, by contrast, are pretrained on small, manually-curated corpora that limit scalability and writing style diversity, a bottleneck e…