AI 中文总结
本研究提出最小循环行为记忆理论,通过兼容关系和熵最小化刻画部分可观测下模仿所需记忆,并设计测量协议,在操作任务中验证了码率需求,同时发现学习该表示的困难与监督退火策略。
AI 中文摘要
在部分可观测条件下,重现指定专家所需的最小循环记忆是多少?瞬时需求是专家行为商的条件熵,但循环记忆还必须保留那些在将来观测恢复之前不会丢失的区分。我们通过一个兼容关系来刻画这种最小循环行为记忆:在传递性条件下,其类别达到精确最小值,而一般情况则是在封闭兼容状态分配上的熵最小化,并在有限实例上给出精确证书。一种单一载体测量协议区分了行为充分性、多余码率以及由观测或其他记忆路径携带的信息;实验比特需求指的是在所声明占用下的诱导符号行为模型。在操作任务中,随着隐藏模式增长到512,学习到的码率保持在接近零比特和两比特需求附近,而预期记忆遵循2→1→0的需求,尽管在等待期间瞬时需求为零。学习这种表示仍然困难:事件无关的未来行为监督产生36/40个充分种子(一个冻结配置),并将最长视野像素设置从0/8提高到6/8个充分保留种子(闭环成功率从0.08提高到0.57)。在未修改的社区基准上,该协议认证了延迟无关的需求,充分的码在中延迟时匹配这些需求。监督有助于承诺,但可能诱发预测盈余;退火该监督使模仿和码率训练减少盈余,将信息论目标与学习能力区分开来。
英文摘要
What is the least recurrent memory needed to reproduce a specified expert under partial observability? The instantaneous requirement is the conditional entropy of the expert's behavioral quotient, but recurrence must also preserve distinctions that future observations will not restore before use. We characterize this minimal recurrent behavioral memory by a compatibility relation: under transitivity its classes attain the exact minimum, while the general case is an entropy minimization over closed compatible state assignments, with exact certificates on finite instances. A sole-carrier measurement protocol separates behavioral sufficiency, excess code rate, and information carried by observations or other memory paths; experimental bit requirements refer to the induced symbolic behavioral model under the stated occupancy. Across manipulation tasks, learned code rates remain near zero- and two-bit requirements as hidden modes grow to $512$, and anticipatory memory follows a $2\to1\to0$ requirement despite zero instantaneous demand during waiting. Learning this representation remains difficult: event-agnostic future-behavior supervision yields $36/40$ sufficient seeds with one frozen configuration and improves the longest-horizon pixel setting from $0/8$ to $6/8$ sufficient held-out seeds (closed-loop success from $0.08$ to $0.57$). On unmodified community benchmarks, the protocol certifies delay-independent requirements, which sufficient codes match at mid-delay. The supervision aids commitment but can induce predictive surplus; annealing it lets imitation and rate training reduce that surplus, separating the information-theoretic target from the ability to learn it.
Comments46 pages, 10 figures. Code: https://github.com/XianyaoLi/DIACRITIC