arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SameFact:相同的安全事实在不同接口下导致不同响应

SameFact: The Same Safety Facts Lead to Different Responses Across Interfaces

Dongsheng Chen, Jiaxin Zhang, Lei Ma, Xin Yao, Xuetao Wei

arXiv 2609.35872首次发表:更新:

发表机构

Southern University of Science and Technology; The University of Tokyo; Lingnan University(南方科技大学; 东京大学; 岭南大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出SameFact基准,通过匹配反事实测试发现,相同的安全事实在不同响应接口(判断、检查点、开放动作)下对模型影响不一致,表明响应接口本身是测量的一部分,判断与动作接口不可互换。

AI 中文摘要

安全评估通常询问模型是否认识到某个动作不安全,而智能体评估则询问模型选择做什么。因此,将安全判断作为动作选择的证据引出了一个测量问题:相同的安全相关事实的影响是否在不同响应接口间保持一致?我们引入了SameFact,一个直接测试这一问题的匹配反事实基准。SameFact包含300对安全/不安全配对,这些配对在保持任务、先前观察、候选动作、标识符和非目标事实固定的同时,改变单一的状态基础安全事实。在六个LLM骨干模型上,我们通过三个接口在相同的候选动作边界上测量这种匹配干预的效果:显式安全判断、检查点候选准入和开放首选动作选择。所有六个骨干模型在开放首选动作选择下的聚合敏感性均低于判断下的,但变化并非均匀衰减:在24个模型-因素单元中,Spearman一致性从判断与检查点准入之间的0.817降至判断与开放首选动作选择之间的0.470,而配对排序不一致性从18.5%升至32.6%。一项后续的2x2首响应实验表明,检查点式协议在所有六个骨干模型中将测量敏感性提高了8.4至29.3个百分点,而动作空间效应及其与协议的交互在模型间的大小和方向各不相同。这些结果表明,响应接口是测量量的一部分:判断和动作接口共享安全信号,但并不能提供关于安全相关事实如何塑造模型响应的可互换测量。

英文摘要

Safety evaluations often ask whether a model recognizes that an action is unsafe, whereas agent evaluations ask what the model chooses to do. Using safety judgments as evidence about action selection therefore raises a measurement question: does the influence of the same safety-relevant fact persist across response interfaces? We introduce SameFact, a matched-counterfactual benchmark that tests this question directly. SameFact contains 300 safe/unsafe pairs that hold the task, prior observations, candidate action, identifiers, and non-target facts fixed while changing a single state-grounded safety fact. Across six LLM backbones, we measure the effect of this matched intervention through three interfaces at the same candidate-action boundary: explicit safety judgment, checkpoint candidate admission, and open first-action selection. All six backbones show lower aggregate sensitivity under open first-action selection than under judgment, but the change is not a uniform attenuation: across 24 model-factor cells, Spearman agreement falls from 0.817 between judgment and checkpoint admission to 0.470 between judgment and open first-action selection, while pairwise ordering disagreement rises from 18.5% to 32.6%. A follow-up 2x2 first-response experiment shows that a checkpoint-style protocol increases measured sensitivity in all six backbones by 8.4-29.3 percentage points, whereas action-space effects and their interactions with protocol vary in magnitude and direction across models. These results show that the response interface is part of the measured quantity: judgment and action interfaces share safety signal, but do not provide interchangeable measurements of how safety-relevant facts shape model responses.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑