Researchers have introduced HUMAID-NER, a new dataset designed for extracting structured information from disaster-related tweets. This dataset includes 60,000 English tweets annotated with ten operational entity types and approximately 175,000 entity spans. To handle the task, a hybrid pipeline combining a spaCy transformer model, domain-specific patterns, and regular expressions was used for annotation. Additionally, a joint multitask learning framework with a RoBERTa-large encoder was developed to perform both named entity recognition and event classification, achieving strong performance on the validation set. AI
IMPACT This dataset and framework could improve the speed and accuracy of information extraction during humanitarian crises.
RANK_REASON The item describes a new dataset and a proposed multitask learning framework for named entity recognition and event classification, published on arXiv. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- CatalyzeX
- DagsHub
- EntityRuler
- Gotit.pub
- Hugging Face
- Humaid Fayez Humaid Khalfan
- HUMAID-NER
- RoBERTa-large
- ScienceCast
- SpaCy
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →