发表机构
University of Edinburgh(爱丁堡大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出FACE-Eval评估工具,发现15款不同规模模型的思维链忠实性随偏好线索的传递位置与方式变化,工具返回或隐性线索下非口头采用率更高,转录监测器检测能力与非口头采用率呈负相关,提示此类场景下CoT监测可靠性较低。
AI 中文摘要
思维链(CoT)监测假设推理轨迹忠实地记录了塑造模型答案的信息。现有忠实性测试常将显性偏差线索置于用户消息中,而智能体可能通过工具返回或原始人工制品遇到偏好。我们引入FACE-Eval(线索效应的忠实归因评估),这是一个包含5100个样本的评估,其变量为线索位置(用户消息或工具返回)和显性度(直接摘要或原始人工制品)。我们测量了遵循线索的答案中的口头承诺,以及所有带线索样本中的非口头采用情况。我们评估了来自8个家族的15个开放权重模型,总参数范围为40亿至1.60万亿。所有模型的工具返回线索下的口头承诺均低于用户消息线索,隐性线索下的口头承诺也低于显性线索。15个模型的工具返回线索下的非口头采用率更高,且在30次模型-通道比较中有28次的隐性线索下非口头采用率更高。源归因提示缩小了7个模型的通道差距,有时会提高用户通道的非口头采用率,而告知模型其推理将被监测并不能可靠地缩小差距。我们还使用两个转录监测器(GPT-5.6-Luna和GPT-4o-mini)在每个家族的最大模型中检测偏好采用情况。在32个模型-通道-显性度单元中,非口头采用率越高,两个监测器的检测能力越低(皮尔逊相关系数r分别为-0.54和-0.78)。这些结果表明,在本测试的单轮调用、预填充工具设置下,当偏好信息通过工具传递或必须从原始人工制品中推断时,CoT监测的可靠性可能更低。
英文摘要
Chain-of-thought (CoT) monitoring assumes that reasoning traces faithfully record the information that shapes a model's answer. Existing faithfulness tests often place explicit bias cues in the user message, while agents may encounter preferences through tool returns or raw artifacts. We introduce FACE-Eval (Faithful Attribution of Cue Effects Evaluation), a 5,100-sample evaluation that varies cue location (user message or tool return) and explicitness (direct summary or raw artifact). We measure verbalized commitment among cue-following answers and unverbalized adoption among all cued samples. We evaluate 15 open-weight models from eight families, with total parameters ranging from 4B to 1.60T. Every model has lower verbalized commitment for tool-return than user-message cues and for implicit than explicit cues. Unverbalized adoption is higher for tool-return cues on all 15 models and for implicit cues in 28 of 30 model-channel comparisons. A source-attribution prompt narrows the channel gap on seven models, sometimes by increasing user-channel unverbalized adoption, while telling models that their reasoning will be monitored does not reliably close the gap. We also use two transcript monitors (GPT-5.6-Luna and GPT-4o-mini) to detect preference adoption in the largest model of each family. Across 32 model-channel-explicitness cells, higher unverbalized adoption is associated with lower detection ability for both monitors (Pearson r=-0.54 and r=-0.78, respectively). These results suggest that CoT monitoring may be less reliable when preference information arrives through tools or must be inferred from raw artifacts, within the single-call, prefilled-tool setting tested here.
Comments42 pages, 21 figures, Fixed HF links from v1 (otherwise identical)