Researchers have developed a formal law to quantify the performance uplift gained from using diverse large language model (LLM) ensembles. This law decomposes ensemble lift into "rescue" and "damage" components, providing a heuristic for predicting performance based on metrics like accuracy-adjusted correctness correlation ($\phi_{\mathrm{adj}}$). The proposed heuristic was tested on over 767,000 inferences across ten open-weight models and multiple benchmarks, demonstrating strong predictive power that transferred effectively to unseen datasets. AI
IMPACT Provides a framework for optimizing LLM ensemble performance by understanding and leveraging diversity.
RANK_REASON The cluster contains two identical arXiv papers detailing a new research finding and methodology.
Read on arXiv cs.MA (Multiagent) →
AI-generated summary · Google Gemini · from 3 sources. How we write summaries →