PulseAugur
EN
LIVE 09:45:15

New audit method detects hidden backdoors in language model explanations

Researchers have developed a new method called Groundedness Drift to audit language model classifiers for hidden backdoors. This technique measures how well an explanation for a model's classification remains consistent with the input data. When tested on two 7B parameter models across various datasets and attack types, Groundedness Drift proved more effective than existing detectors in identifying these backdoors while maintaining a low false positive rate. An additional evaluation using Unsupported Groundedness was also performed to assess performance under more complex camouflage scenarios. AI

IMPACT Introduces a novel auditing technique that could improve the safety and reliability of language model classifiers.

RANK_REASON Academic paper detailing a new method for auditing language models. [lever_c_demoted from research: ic=1 ai=1.0]

Read on arXiv stat.ML →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

New audit method detects hidden backdoors in language model explanations

COVERAGE [1]

  1. arXiv stat.ML TIER_1 English(EN) · Yang Liu, Ran Zou ·

    When Explanations Betray Backdoors: Black-Box Auditing for Language Model Classifiers

    arXiv:2608.12623v1 Announce Type: cross Abstract: Language model classifiers with explanations are used for moderation, routing, topic triage, and low-resource annotation. We study black-box auditing when the defender has only clean calibration data without trigger information bu…