PulseAugur
实时 06:36:41
English(EN) En-ViMedNER: An English-Vietnamese Parallel Biomedical Corpus with UMLS Semantic Type Annotations

新的 En-ViMedNER 语料库助力越南语生物医学 NLP 研究

研究人员推出了 En-ViMedNER,这是首个英越平行生物医学命名实体识别 (NER) 语料库。该语料库使用统一医学语言系统 (UMLS) 语义类型进行标注,为与现有基于英语的资源进行直接比较提供了一个共享的跨语言标签空间。En-ViMedNER 包含超过 4,300 对 PubMed 摘要和近 203,000 个对齐的实体提及对,通过自动翻译、专家编辑和 LLM 辅助方法相结合构建而成。该语料库已针对越南语输入和跨语言 NER 任务进行了评估,基线模型的 F1 分数最高可达 53.78。 AI

影响 支持为越南语开发生物医学 NLP 工具,改进医疗保健 AI 应用和医学信息提取。

排序理由 该条目描述了一篇介绍用于 NLP 研究的新颖数据集的学术论文。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.CL 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

新的 En-ViMedNER 语料库助力越南语生物医学 NLP 研究

本文如何被排名

Signal score
29 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
该条目描述了一篇介绍用于 NLP 研究的新颖数据集的学术论文。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, other
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

完整方法见我们的编辑标准

报道来源 [1]

  1. arXiv cs.CL TIER_1 English(EN) · Nhu Vo, Phuong Nguyen, Nu Uyen Phuong Le, Inigo Jauregi Unanue, Dung D. Le, Massimo Piccardi, Wray Buntine ·

    En-ViMedNER:一个带有UMLS语义类型标注的英越平行生物医学语料库

    arXiv:2608.29890v1 Announce Type: new Abstract: Biomedical Named Entity Recognition (NER) is fundamental to healthcare AI applications, including clinical decision support and medical information extraction. While corpora with Unified Medical Language System (UMLS) annotations, s…