Researchers have developed BLOOM-WILT, a novel auditing pipeline designed to elicit rare behaviors in large language models (LLMs) during deployment. This method uses a full auditing pipeline that learns from scored interactions and adaptively reweights the target model's decoding process to favor behavior-relevant generations. Evaluations across four target models and eight behaviors demonstrated that BLOOM-WILT significantly outperforms baseline auditors, increasing the presence of rare behaviors from 51% to 100% in some cases, without compromising output probability. AI
IMPACT This new auditing technique could improve the safety and reliability of deployed LLMs by surfacing rare failure modes.
RANK_REASON This is a research paper detailing a new method for auditing LLMs. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →