arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

奖励何时能引导对状态的理解?一种隐藏自动机工具与群语言边界

When Does Reward Teach State? A Hidden-Automaton Instrument and a Group-Language Warning Signal

James E. Allchin

arXiv 2607.11953首次发表:更新:

AI 中文总结

研究强化学习中高奖励与潜在状态理解的关系,通过将任务表示为隐藏自动机,利用三个可控轴(优化器强度、任务结构、观察信息量)来分别测量奖励成功和潜在状态学习,区分感知差距和规划差距,得出高奖励非任务理解证据且智能体恢复潜在状态可预测的结论。

AI 中文摘要

强化学习中,获得高奖励的智能体是代表了任务的潜在状态,还是只是与奖励相关的捷径?由于“真实状态”未定义,该问题通常无法回答。本文通过一个白盒工具使其可精确回答:将任务表示为隐藏确定性有限自动机(DFA),让智能体观察符号流并在部分控制下间歇性选择下一个符号,接受时给予一个稀疏终端奖励。知道自动机可免费获得最优回报(奖励成为可解释的归一化分数)和每一步的精确潜在状态。奖励成功和潜在状态学习成为可分别测量的量,其耦合由三个可控轴控制。包括优化器强度,如弱策略强化学习下智能体通过状态探测获得奖励,不同架构有差异;任务结构,排列(群语言)结构可在训练前从转移函数计算得出,能标记感知差距;观察信息量,无标签辅助在观察无状态时无用,随观察揭示状态的程度恢复状态。结果是区分了仅基于奖励评估无法区分的感知差距(潜在状态虽可表示但不能线性恢复)和规划差距(状态可恢复但未使用)。高奖励并非任务理解的证据,智能体是否恢复潜在状态可提前预测。

英文摘要

Does a reinforcement-learning agent that earns reward learn its task's hidden state? We study this question with hidden finite automata that the agent partially controls. Because each automaton is known, we can normalize reward by the best achievable return and probe the network for the true state at every step. Together the two measurements separate failures that reward alone conflates. An agent can encode too little of a state its network could hold, or encode the state and still control poorly. Weak on-policy RL matches random play while the state probe stays at chance. State learning depends on the optimizer, the training budget, and the task's structure. Permutation automata provide a warning before training: no input symbol maps two distinct states to the same successor. On a stratified held-out set, 86 of 103 permutation automata fail the state probe, and the classification is stable across probe read-outs and recovery thresholds. Most of these failures come with weak reward. High reward without the state occurs but is rare. Non-permutation automata can also fail. Oracle-normalized reward alone therefore does not establish that the task's state was learned.

Comments22 pages, 9 figures, 10 tables (12-page main text). Ancillary file hidden-automata-rl-code.zip holds reproduction code and the per-run data behind every table and figure. v4: held-out numbers from the 5-seed uniform-budget rerun (86/103), observation-only control, 5-seed intervention, recovery-threshold and minimality robustness, experimental protocol in the main text, claim scoping

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑