Irt
PulseAugur coverage of Irt — every cluster mentioning Irt across labs, papers, and developer communities, ranked by signal.
2 day(s) with sentiment data
-
New IRT model predicts LLM abilities efficiently with 85% cost reduction
Researchers have developed a cost-efficient method for evaluating large language models (LLMs) by predicting their performance on unseen tasks. This approach utilizes a modified multidimensional item response theory (IR…
-
Study finds HLE benchmark measures general reasoning, not distinct LLM capabilities
A new study analyzing the Humanity's Last Exam (HLE) benchmark has found that its multiple-choice subset, comprising 428 items, primarily measures a single general reasoning factor rather than distinct subject-domain ca…
-
Researchers explore diminishing returns in LLM benchmark size using IRT
Researchers explored the diminishing returns of increasing benchmark size for Large Language Models (LLMs) using Item Response Theory (IRT). They found that while IRT provides a theoretical framework for measuring the i…
-
New Framework for Evaluating RAG Systems by Question Granularity
Researchers have introduced HieraRAG, a hierarchical framework for evaluating retrieval-augmented generation (RAG) systems by analyzing question granularity. This framework aims to help practitioners determine the optim…
-
Researchers quantify and mitigate socially desirable responding in LLMs
Researchers have developed a new framework to identify and reduce socially desirable responding (SDR) in large language models (LLMs) when they are evaluated using self-report questionnaires. This SDR, where models prov…