Researchers have introduced ArabicDialectSafety, a new dataset designed to evaluate the safety of Arabic content across various dialects. The dataset includes 25,071 prompts in six Arabic varieties and is annotated for dialect and seven harm categories. Fine-tuned MARBERTv2 demonstrated superior performance in safety classification compared to large language models, achieving high Macro-F1 scores. The study also found that integrating dialect information at the representation level is most effective, though performance lags for less-resourced Maghrebi dialects. AI
IMPACT This benchmark will enable more nuanced safety evaluations for LLMs handling diverse Arabic dialects.
RANK_REASON The cluster contains a research paper introducing a new benchmark dataset for a specific NLP task. [lever_c_demoted from research: ic=1 ai=1.0]
- ArabicDialectSafety
- arXiv
- Egyptian
- MARBERTv2
- Modern Standard Arabic
- North Africans
- Palestinian
- Syria
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →