A new study published on arXiv evaluates the performance of sentiment analysis tools and large language models (LLMs) on social media texts. The research found that even human annotators exhibit only fair agreement when classifying sentiment, highlighting the inherent subjectivity of the task. Among the evaluated tools, Twitter-roBERTa-base demonstrated the strongest alignment with human ratings, particularly for binary sentiment classification, while LLMs like Qwen3-32B, GPT-OSS-120B, and Llama-4-Maverick-17B showed moderate to substantial agreement with humans and strong agreement among themselves. AI
IMPACT Highlights the need for domain-specific fine-tuning and human-centered evaluation for reliable social media sentiment analysis.
RANK_REASON Academic paper evaluating LLM and tool performance on a specific task. [lever_c_demoted from research: ic=1 ai=1.0]
- Cohen's kappa
- Fleiss' kappa
- GPT-OSS 120B
- Himarsha R Jayanetti
- Llama-4-Maverick-17B
- Qwen3 32B
- TextBlob
- Twitter-roBERTa-base
- VADER
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →