Researchers have developed SLAyiNG, a novel dataset comprising over 500 queer slang terms validated by the community, aimed at improving Natural Language Processing (NLP) systems' understanding of queer vernacular. The dataset addresses the issue of NLP models misclassifying queer language or generating negative responses. Initial findings using SLAyiNG indicate that while some language models are not biased against the queer community, they still struggle to process its language, highlighting a distinction between representation bias and linguistic bias. Furthermore, the dataset reveals that model performance on queer slang varies across different queer subcommunities, with notably poorer performance for slang associated with African-American and Latine communities. AI
IMPACT This dataset could improve the performance and fairness of AI systems for LGBTQ+ users.
RANK_REASON The item is an academic paper detailing a new dataset and research findings. [lever_c_demoted from research: ic=1 ai=1.0]
- arXiv
- Hugging Face
- language models
- Latine communities
- Leonor Veloso
- natural language processing
- NLP systems
- queer slang
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →