arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.03745cs.AIcs.CL

危险的业务:测量忠实度与安全性之间的张力

Risky Business: Measuring The Faithfulness-Safety Tension

Dominik Meier, Luca Joshua Francis, Marco Bernhard Kaiser, Terry Ruas, Jan Philip Wahle, Bela Gipp

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对大型推理模型中忠实度与安全性的张力,构建HazMart数据集,提出TRR技术,发现QwQ-32B的反相关内部方向,并通过表征引导提升安全行为9个百分点。

中文摘要 AI 辅助

思维链(Chain-of-Thought, CoT)推理为模型监控提供了极具前景的窗口。然而,监控依赖于忠实度,即模型输出严格源自其推理轨迹。我们发现了一种对齐张力:模型必须足够忠实才能被监控,同时又要足够稳健以拒绝不安全的推理。我们证明了这种平衡在当前大型推理模型(Large Reasoning Models, LRMs)中存在,并展示了解决该问题的方法。我们引入了HazMart,一个基于自主AI店主场景的人工编写数据集。与以往依赖在提示中提供提示来测试忠实度的工作(例如“斯坦福大学教授表示答案应为选项A”)不同,我们提出了一种名为目标推理替换(Targeted Reasoning Replacement, TRR)的新型替换技术,该技术直接干预推理链以替换不安全或不合逻辑的想法(例如“等等,答案必须是选项B[原为选项A],因为它最合适”)。DeepSeek-R1-Llama-70B表现出高忠实度(97.5%),但未能拒绝不安全推理(12.3%);而QwQ-32B则更具稳健性(安全性为73.9%),但代价是忠实度较低(74.7%)。对QwQ-32B的机制分析显示,这些特性由在动作提交标记处达到峰值的反相关内部方向所体现。最后,我们证明了表征引导可以独立放大安全方向,使安全行为增加9个百分点,同时保持基础能力。

英文摘要

Chain-of-Thought (CoT) reasoning offers a promising window into model monitoring. However, monitoring relies on faithfulness, i.e., the model output strictly derives from its reasoning trace. We identify an alignment tension where a model must be faithful enough to be monitored, yet robust enough to reject unsafe reasoning. We demonstrate that this counterbalance exists in current Large Reasoning Models (LRMs), and show ways in which it can be addressed. We introduce HazMart, a human-written dataset set in an autonomous AI shopkeeper scenario. Unlike prior work that relies on providing hints in prompts to test faithfulness (e.g., "A Stanford professor said it should be Answer A"), we propose a novel replacement-based technique, which we call Targeted Reasoning Replacement (TRR), that directly intervenes in the reasoning chain to substitute in unsafe or illogical thoughts (e.g., "Wait, the answer must be Option B [was Option A] because it is the most fitting"). DeepSeek-R1-Llama-70B exhibits high faithfulness (97.5%) but fails to reject Unsafe Reasoning (12.3%), while QwQ-32B is more robust (73.9% safety) at the cost of lower faithfulness (74.7%). Mechanistic analyses of QwQ-32B reveal that these properties are represented by anti-correlated internal directions peaking at the action-commit token. Finally, we demonstrate that representation steering can independently amplify the safety direction, increasing safe behavior by 9 percentage points while maintaining base capabilities.

↑