Researchers have introduced En-ViMedNER, the first parallel English-Vietnamese biomedical Named Entity Recognition (NER) corpus. This corpus is annotated with Unified Medical Language System (UMLS) semantic types, providing a shared cross-lingual label space for direct comparison with existing English-based resources. En-ViMedNER contains over 4,300 PubMed abstract pairs and nearly 203,000 aligned entity-mention pairs, constructed using a combination of automatic translation, expert editing, and LLM-assisted methods. The corpus has been evaluated for both Vietnamese-input and cross-lingual NER tasks, with baseline models achieving F1 scores up to 53.78. AI
IMPACT Enables development of biomedical NLP tools for Vietnamese, improving healthcare AI applications and medical information extraction.
RANK_REASON The item describes a new academic paper introducing a novel dataset for NLP research. [lever_c_demoted from research: ic=1 ai=1.0]
- English
- En-ViMedNER
- MedMentions
- MedMentions ST21pv
- natural language processing
- PubMed
- Unified Medical Language System
- Vietnamese
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →