A new study published on arXiv explores the reliability of using large language models (LLMs) for cultural analytics, specifically examining their ability to annotate English song lyrics for social constructs. The research evaluated five LLMs on their consistency in measuring self-esteem, self-control, seeking belonging, and seeking recognition. Findings indicate that LLM reliability varies significantly by construct, with self-esteem showing the most stable measurements and seeking recognition being less consistent. The study suggests that while LLM-generated labels contain usable signals for downstream classification, reporting on repeated-measurement stability and cross-model convergence is crucial before these annotations are accepted as scalable measurements in cultural analytics. AI
IMPACT Highlights the need for rigorous validation of LLM outputs in cultural analytics, impacting how researchers use AI for text annotation.
RANK_REASON The cluster contains a research paper published on arXiv detailing a study on LLM capabilities. [lever_c_demoted from research: ic=1 ai=1.0]
- arXiv
- English song lyrics
- Hugging Face
- large language models
- seeking belonging
- Self Control
- Self Esteem
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →