当解释背叛后门:语言模型分类器的黑盒审计
When Explanations Betray Backdoors: Black-Box Auditing for Language Model Classifiers
- University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校)
- University of California, Irvine(加利福尼亚大学欧文分校)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文针对仅拥有干净校准数据的黑盒场景,提出接地漂移方法,在5%干净FPR预算下,于7B主干等实验中实现了比现有检测器更优的后门检测性能。
AI中文摘要:
带有解释的语言模型分类器被用于内容审核、路由、主题分类及低资源标注。我们研究防御者仅拥有干净校准数据(无触发器信息)但可向分类器查询标签及简短理由或引用证据时的黑盒审计。我们引入「接地漂移(Groundedness Drift)」,这是衡量答案摘要是否仍基于输入的轻量分数。在两个7B主干、五个数据集及四种常见非自适应OpenBackdoor式攻击族下,当干净假正例率(clean-FPR)预算为名义值5%时,接地漂移在所有情况下均优于所有对比检测器,取得更高的AUROC和更低的残留目标攻击成功率(ASR)。随后我们评估「未支持接地(Unsupported Groundedness)」——一种针对解释伪装压力场景的多探针升级方案,该方案增强了信号但未弥合自适应差距。
英文摘要:
Language model classifiers with explanations are used for moderation, routing, topic triage, and low-resource annotation. We study black-box auditing when the defender has only clean calibration data without trigger information but can ask the classifier for a label plus a short rationale or quoted evidence. We introduce Groundedness Drift, a lightweight score measuring whether the answer summary remains grounded in the input. Across two 7B backbones, five datasets, and four common non-adaptive OpenBackdoor-style attack families, Groundedness Drift achieves higher AUROC and lower residual target ASR than every compared detector in all cases at a nominal 5\% clean-FPR budget. We then evaluate Unsupported Groundedness, a multi-probe escalation for explanation-camouflage stress cases. Unsupported Groundedness improves signals but does not close the adaptive gap.