Researchers have developed a new benchmark, Principle-Bench, to evaluate the trustworthiness of Large Language Models (LLMs) when used as judges in principle-based regulation. The benchmark assesses LLMs across four axes: accuracy, paraphrase robustness, adversarial robustness, and calibration, using 168 cryptoasset financial-promotion scenarios mapped to UK FCA principles. Findings indicate that even large LLMs struggle with adversarial inputs, demonstrating significant drops in accuracy and poor agreement with other models when faced with keyword-stuffed or perturbed data, highlighting the risks of "compliance theatre." AI
IMPACT Highlights the need for robust evaluation of LLMs in regulatory contexts, particularly against adversarial attacks, to prevent 'compliance theatre'.
RANK_REASON The cluster contains an academic paper introducing a new benchmark and methodology for evaluating LLMs. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →