Linear Probes Detect Task Format, Not Reasoning Mode in Language Model Hidden States
线性探针检测语言模型隐藏状态中的任务格式,而非推理模式
机构 * Horizon Research(远景研究) ; Meta ; Apple(苹果公司) ; Northeastern University(东北大学)
AI总结 通过线性探针实验发现,大语言模型隐藏状态中看似分离的推理模式实际上由任务格式(如来源、选项数、响应长度)混淆导致,而非真正的推理计算结构。
Comments Accepted in the 6th Workshop on Trustworthy NLP, ACL 2026