A new resource called AtlasNLP has been developed to map the geographic representation within natural language processing (NLP) datasets. This atlas, which includes a human-curated set (AtlasNLP-Gold) and a large collection derived from the Association for Computational Linguistics (AtlasNLP-Core), tracks both the populations represented and the production locations of over 13,000 NLP dataset records. Initial findings indicate significant disparities in dataset coverage across countries and tasks, highlighting an asymmetry between where datasets are produced and the populations they represent. The research emphasizes that language coverage does not equate to geographic representation, revealing critical gaps in current dataset documentation and advocating for more explicit geographic metadata. AI
IMPACT Highlights critical gaps in NLP dataset documentation, potentially guiding future data collection and evaluation efforts for more equitable AI.
RANK_REASON The cluster contains a research paper detailing a new resource and findings in NLP dataset representation. [lever_c_demoted from research: ic=1 ai=1.0]
- arXiv
- Association for Computational Linguistics
- AtlasNLP
- AtlasNLP-Core
- AtlasNLP-Gold
- Hugging Face
- natural language processing
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →