arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

合规检测器读取什么?激活探针与防护模型的审计

What Do Compliance Detectors Read? An Audit of Activation Probes and Guard Models

Saisab Sadhu, Aadit Sengupta, Vinay Kumar Sankarapu, Pratinav Seth

arXiv 2608.16852首次发表:更新:

发表机构

Lexsi Labs(雷克西实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究通过审计合规检测器发现其存在规则盲视问题,引入无需训练的ICS检测器,发布相关基准以推动规则盲视的进一步研究。

AI 中文摘要

部署的语言模型中的监管合规监测正日益作为法律和审计控制手段,用于对照涵盖数据保护、医疗保健、金融监管及平台政策的书面规则检查模型输出。只有当检测器的判断取决于所述规则而非场景的表面特征时,此类监测才有意义。我们表明,当前所有合规检测器类别均不满足这一条件,我们将这种失败称为“规则盲视”。对于我们测试的每个防护模型和激活探针,包括一个能正确引用管辖条款的策略条件防护模型,当该条款被其宽松对应条款替换时,其判断几乎没有变化;删除、置换或替换管辖规则后,检测准确率保持不变。我们构建了一个专用基准,将两条规则与两个场景交叉,使得单独的规则或场景均无法预测标签,该基准确认了在现有基准未排除的设计下存在规则盲视问题,且我们测试的快速检测器均无法解决该问题,只有逐步推理能规避它。大规模审计需要无需重新训练的检测器,因此我们引入了内部合规评分(Internal Compliance Score, ICS):一种无需训练的激活读出器,由十对标注样本校准,通过单一投影评分。我们对ICS进行了与所审计防护模型相同的审查:其未达到击败简单基线的预先注册标准,且词袋模型与其总体泛化能力完全匹配。ICS仍有用,因为它成本低廉,可用于审计四个部署的防护模型、一个8B零样本评判器及十三个基准,且当用于对候选响应排名时,可提高机械验证的通过率,不过自适应白盒攻击会消除这一增益。我们发布了反事实协议和交叉规则基准,以便未来能在探针和防护模型的相关主张中测试规则盲视问题。

英文摘要

Regulatory compliance monitoring in deployed language models is increasingly implemented as a legal and audit control, checking model outputs against written rules spanning data protection, healthcare, financial regulation, and platform policy. Such monitoring is meaningful only if a detector's verdict depends on the stated rule rather than on surface features of the scenario. We show this condition fails across the current class of compliance detectors, a failure we call rule blindness. Deleting, permuting, or substituting the governing rule leaves detection accuracy unchanged for every guard and activation probe we test, including a policy-conditioned guard that correctly cites the governing clause yet barely changes its verdict when that clause is swapped for its permissive counterpart. A purpose-built benchmark crossing two rules with two scenarios, so that neither alone predicts the label, confirms the failure under a design no prior benchmark rules out, and shows that step by step reasoning, not any fast detector we test, is what escapes it. Auditing at scale requires a retraining-free detector, so we introduce the Internal Compliance Score (ICS): a training-free activation readout calibrated from ten labelled pairs and scored by a single projection. We hold ICS to the same scrutiny as the guards it audits: a pre-registered criterion for beating trivial baselines is not met, and a bag-of-words model matches its pooled generalisation exactly. It remains useful because it is inexpensive, letting us audit four deployed guard models, an 8B zero-shot judge, and thirteen benchmarks, and it raises the mechanically verified pass rate when used to rank candidate responses, though an adaptive white-box attack removes this gain. We release the counterfactual protocol and crossed-rule benchmark so rule blindness can be tested in future probe and guard claims.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑