PulseAugur
EN
LIVE 10:48:45

New dataset SLAyiNG aims to improve NLP for queer slang

Researchers have developed SLAyiNG, a novel dataset comprising over 500 queer slang terms validated by the community, aimed at improving Natural Language Processing (NLP) systems' understanding of queer vernacular. The dataset addresses the issue of NLP models misclassifying queer language or generating negative responses. Initial findings using SLAyiNG indicate that while some language models are not biased against the queer community, they still struggle to process its language, highlighting a distinction between representation bias and linguistic bias. Furthermore, the dataset reveals that model performance on queer slang varies across different queer subcommunities, with notably poorer performance for slang associated with African-American and Latine communities. AI

IMPACT This dataset could improve the performance and fairness of AI systems for LGBTQ+ users.

RANK_REASON The item is an academic paper detailing a new dataset and research findings. [lever_c_demoted from research: ic=1 ai=1.0]

Read on arXiv cs.CL →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

New dataset SLAyiNG aims to improve NLP for queer slang

COVERAGE [1]

  1. arXiv cs.CL TIER_1 English(EN) · Leonor Veloso, Lea Hirlimann, Lucija Mihi\'c Zidar, Philipp Wicke, Valentin Hofmann, Hinrich Sch\"utze ·

    SLAyiNG: A Diverse and Community-validated Dataset of Queer Slang

    arXiv:2509.17449v2 Announce Type: replace Abstract: Queer vernacular is rarely studied in NLP, despite advancements in resources and evaluation for other sociolects and informal language. Because of this, NLP systems often process queer language incorrectly, e.g., they misclassif…