A new paper explores the effectiveness of Large Language Models (LLMs) in evaluating explanations generated for recommender systems. Researchers found that while LLMs can mimic human rating patterns and show moderate correlation with human judgments, their absolute agreement is low and inconsistent across different model sizes and evaluation criteria. The study suggests practical recommendations for using LLMs in this capacity, including using concise prompts, favoring larger models, pre-testing evaluation constructs, and auditing explanations for factual accuracy. AI
IMPACT LLMs show potential but also limitations in evaluating recommender system explanations, suggesting careful implementation is needed.
RANK_REASON Academic paper on the utility of LLMs for evaluating recommender system explanations. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →