Researchers have developed NormasTCU, a new dataset for Brazilian Portuguese Information Retrieval (IR) that includes 14,469 legal documents and human relevance judgments. The study evaluated the effectiveness of using Large Language Models (LLMs) as judges for relevance assessment in this specialized domain. While LLMs demonstrated a positive scoring bias and only fair to moderate agreement with human judgments, they often produced system rankings highly correlated with human assessments for metrics like nDCG@10 and MRR. AI
IMPACT LLMs can potentially scale relevance assessment in specialized, non-English domains, improving IR system evaluation efficiency.
RANK_REASON The cluster describes a new dataset and an evaluation of LLM performance on a specific task within a research paper. [lever_c_demoted from research: ic=1 ai=1.0]
Read on arXiv cs.IR (Information Retrieval) →
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →