Researchers have developed Meddies-PII, a multilingual framework designed for extracting personally identifiable information (PII) from clinical documents. This framework includes a large synthetic dataset of one million clinical documents in seventeen languages, generated using attribute-conditioned prompts and validated through consistency checks. The associated Meddies-PII-Model, a BIOES token classifier, demonstrated superior performance on PII extraction benchmarks, achieving a mean F1 score of 0.827, significantly outperforming the strongest baseline at 0.658. The dataset, model, and associated code will be made publicly available to advance research in multilingual clinical de-identification. AI
IMPACT Enhances the accuracy and efficiency of de-identifying sensitive clinical data across multiple languages.
RANK_REASON The cluster describes a new research paper detailing a framework, dataset, and model for a specific NLP task. [lever_c_demoted from research: ic=1 ai=1.0]
- arXiv
- BioEssays
- clinical de-identification
- Hugging Face
- Meddies-PII
- Meddies-PII-Dataset
- Meddies-PII-Model
- personally identifiable information
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →