arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.22850cs.LGcs.AI

测试LLM智能体中功能效价轴的构念效度

Same Outcome, Different Readout: What Does a Steerable Valence Direction in LLMs Represent?

  • The University of Tokyo(东京大学)
  • Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳))
  • The Hong Kong University of Science and Technology(香港科技大学)

机构由 AI 辅助整理,请以论文原文为准。

Weihan Li, Xinlei Chen, Yuhan Song, Xiaofeng Lin, Tianshi Zheng

AI总结:

本研究通过迷宫任务中的受控干预测试LLM智能体中好—坏结果方向的功能效价构念效度,发现该方向具有价值相关功能但非历史不变的标量效价状态。

AI中文摘要:

对比激活方向通常根据它们解码出的内容或它们引导行为的强度来解释。但是,什么证据足以确定这样一个方向所代表的构念,而非用于提取它的对比中的相关特征?我们通过迷宫任务中的好—坏结果方向来研究这个问题,使用受控干预将已实现的结果与得知该结果的信息历史分开。在多个LLM检查点中,基于一种显式结果编码拟合的方向能很好地迁移到另一种编码,表明读出并非绑定于表面形式。相反,当相同的已实现结果通过已宣布和未宣布的历史达到时,迁移显著下降:即使在两种历史都接收到相同的显式结果后,事件后的读出仍强烈依赖于先前的宣布。在匹配的迷宫强化学习运行中,强化学习后的方向对参考MDP剩余回报的预测性显著增强,且在测试位置上策略对该方向的依赖性增强,而这种历史依赖性持续存在。这些结果支持该方向的功能性、价值相关解释,但不支持将其识别为历史不变的标量效价状态。

英文摘要:

Claims about what an internal direction in an LLM represents need evidential constraints beyond an observer's prior beliefs about the system. Decodability and successful activation steering do not, by themselves, establish which construct the direction tracks. This gap is especially consequential for welfare-relevant interpretations, where a proposed functional state must be distinguished from correlated features of the extraction contrast. We treat the question as one of construct validity and study a good-bad outcome direction in a maze task, using controlled interventions that separate the realised outcome from the informational history through which it became known. Across multiple LLM checkpoints, directions fitted on one explicit outcome encoding transfer well to another, indicating that the readout is not tied to surface form. When the same realised outcome is reached through announced and unannounced histories, however, transfer degrades substantially: even after both histories receive the same explicit outcome, the post-event readout remains strongly conditioned on the earlier announcement. In the base model, steering along the direction changes actions, yet removing it leaves the natural cue effect almost intact. In a matched maze-RL run, the post-RL direction becomes substantially more predictive of reference-MDP remaining return and the policy becomes more dependent on it at the tested sites, while the history dependence persists. These dissociations support a functional, value-related interpretation of the direction, but not its identification with a history-invariant scalar valence state.

补充信息

↑