Researchers have developed a new method called Groundedness Drift to audit language model classifiers for hidden backdoors. This technique measures how well an explanation for a model's classification remains consistent with the input data. When tested on two 7B parameter models across various datasets and attack types, Groundedness Drift proved more effective than existing detectors in identifying these backdoors while maintaining a low false positive rate. An additional evaluation using Unsupported Groundedness was also performed to assess performance under more complex camouflage scenarios. AI
IMPACT Introduces a novel auditing technique that could improve the safety and reliability of language model classifiers.
RANK_REASON Academic paper detailing a new method for auditing language models. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →