Researchers have developed a new framework to audit decision systems that exhibit the Rashomon effect, a phenomenon where multiple accurate models produce different predictions. This framework combines ensemble margin with local prediction variability to identify incorrect ensemble predictions. Experiments using transformer models for natural language understanding and fine-tuned large language models for tabular data show that this ensembling approach significantly reduces the risk of unchecked incorrect predictions while only moderately increasing the number of instances requiring human review. AI
IMPACT Introduces a more reliable method for capturing predictive multiplicity in AI systems, potentially improving the safety and trustworthiness of deployed models.
RANK_REASON Academic paper detailing a new auditing framework for machine learning models. [lever_c_demoted from research: ic=1 ai=1.0]
- ensemble margin
- large language models
- local prediction variability
- natural language understanding
- Rashomon effect
- tabular data classification
- transformer models
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →