AI 中文总结
本研究通过环境变化测试检验语言模型反事实报告的可识别性,发现错误来源演示影响报告,而机制绑定可减少影响,强调基准需包含不变性测试。
AI 中文摘要
语言模型的自我报告是关于提示环境中行为的证据,本身并不是自我模型的证据。我们研究了在激活干预下关于情感类状态的反事实报告,并询问当演示环境改变时,报告是否仍然与指定的干预绑定。在三个开放指令模型中,错误来源的演示将报告拉向来源答案族,而明确的机制绑定减少了这种拉力。自我报告基准应在固定干预下包含环境变化不变性测试,然后再将准确性视为自主报告机制的证据。
英文摘要
Language-model self-reports are evidence about behavior in a prompt environment, not by themselves evidence of a self-model. We investigate counterfactual reports about affect-like states under activation interventions and ask whether the report remains bound to the named intervention when the demonstration environment changes. Across three open instruction models, wrong-source demonstrations move reports toward the source answer family, while explicit mechanism binding reduces this pull. Self-report benchmarks should include environment-shift invariance tests under fixed intervention before treating accuracy as evidence for an autonomous report mechanism.