发表机构
Lexsi Labs(Lexsi实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究通过测试五个模型在20个监管及平台政策领域的规则扰动下的裁决变化,发现大语言模型合规系统的裁决对规则扰动常具不变性,防护模型规则敏感性最低且准确性仅略高于随机水平,提示仅靠准确性无法证明裁决基于规则。
AI 中文摘要
大语言模型合规系统的部署基于一个假设:裁决取决于其所给定的监管规则。我们针对五个模型和20个监管及平台政策领域直接验证了这一点:在保持案例固定的情况下,删除、替换或否定管辖规则,检查裁决是否发生变化(OCS),或模型的合规性内部表示是否发生任何变化(ICS-delta)。两者变化都不大:模型的裁决通常对所提供规则的重大扰动具有不变性,且在此处根据其原生分类法进行自定义规则适配评估的防护模型,是五个模型中规则敏感性最低、准确性最差的,仅略高于随机水平(51%,而通用模型为90%-92%)。这反映的是简单案例,而非全面忽视:在删除规则会改变模型先前正确预测的案例中,模型确实会紧密跟踪规则。无论是更好的提示还是对模型内部表示的直接干预,都无法缩小这一差距。仅靠准确性并不能证明合规裁决是基于所提供的规则。
英文摘要
Large language model compliance systems are deployed on the assumption that a verdict depends on the regulatory rule it is given. We test this directly across five models and 20 regulatory and platform-policy domains: delete, swap, or negate the governing rule while holding the case fixed, and check whether the verdict changes (OCS) or the model's internal representation of compliance shifts at all (ICS-delta). Neither moves much: models' verdicts are often invariant to substantial perturbations of the supplied rule, and the guard model, evaluated here under a custom-rule adaptation of its native taxonomy, is the least rule-sensitive and least accurate of the five, barely above chance (51%, versus 90-92% for general-purpose models). This reflects easy cases more than blanket neglect: on cases where deleting the rule changes a previously correct model prediction, models do track it closely. Neither better prompting nor direct intervention on the model's internal representations closes this gap. Accuracy alone does not establish that a compliance verdict is grounded in the supplied rule.