A new study published on arXiv compares the performance of GPT-5.4 against human coders in analyzing open-ended survey responses for inductive content analysis. The research found that GPT-5.4 achieved an Adjusted Rand Index (ARI) of 0.61 for coding and 0.54 for theme generation, which was comparable to the internal consistency among human coders (ARI=0.68) and GPT-5.4 itself (ARI=0.76). These results suggest that LLMs like GPT-5.4 can serve as a scalable tool to support qualitative analysis, particularly at the coding level, though agreement varied across different survey variables. AI
IMPACT LLMs can serve as a scalable tool to support qualitative analysis, particularly at the coding level.
RANK_REASON Research paper published on arXiv comparing LLM performance to human analysis. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →