arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

上下文内神经反馈:大语言模型能否通过特权访问控制其内部表征?

In-Context Neurofeedback: Can LLMs Control Their Internal Representations through Privileged Access?

Koshiro Aoki, Ryota Takatsuki, Gouki Minegishi, Yusuke Haruki, Daisuke Kawahara

arXiv 2609.00904首次发表:更新:

发表机构

Waseda University; University of Sussex; The University of Tokyo(早稻田大学; 萨塞克斯大学; 东京大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对LLMs的神经反馈范式重新设计,发现严格设定下模型无法可靠控制特权内部表征,说明此前相关结论可能依赖表面机制,强调评估LLMs元认知需用特权访问类方法。

AI 中文摘要

大语言模型(LLMs)能否控制自身内部表征对机器元认知和AI安全都至关重要。一项近期研究将神经反馈应用于LLMs,声称它们能控制自身内部表征,但该研究报告的控制可能依赖表面机制而非真正的内部访问,因为其中的控制目标并非特权目标,即第三方可从提示中推断这些目标。我们为LLMs重新设计了神经反馈范式,使控制目标满足特权访问要求,更接近人类认知神经科学中的神经反馈实验。在这一更严格的设定下,模型未展现出对特权内部表征的可靠控制,表明此前报告的控制无法排除其依赖表面机制的可能性。我们的结果表明,对LLMs元认知的严格评估需要依赖特权访问的评估方法。

英文摘要

Whether large language models (LLMs) can control their own internal representations matters for both machine metacognition and AI safety. A recent study applied neurofeedback to LLMs and claimed that they can control their internal representations. However, the reported control may rely on superficial mechanisms rather than genuine internal access because the control targets in that study are not privileged, meaning that a third party can infer them from the prompt. We redesign the neurofeedback paradigm for LLMs so that the control target satisfies the privileged access requirement, which is closer to neurofeedback experiments in human cognitive neuroscience. Under this stricter setting, the models do not demonstrate reliable control over privileged internal representations, suggesting that previously reported control cannot exclude the possibility that it relies on superficial mechanisms. Our results indicate that rigorous assessments of metacognition in LLMs require evaluation methods that demand privileged access.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑