强化学习中的策略复杂度、反应时间与有界理性
Policy Complexity, Reaction Time, and Bounded Rationality in Reinforcement Learning
查看机构详情
- Rensselaer Polytechnic Institute(伦斯勒理工学院)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
本文提出MI-SARSA算法,通过互信息正则化在强化学习中引入认知约束,实现策略复杂度与反应时间的权衡,并揭示鲁棒性与容量的权衡,为有界理性下的序列学习提供模型。
中文摘要 AI 辅助
生物智能体并非在无限计算条件下进行学习。对于人类而言,学习与选择受到感知、注意力和工作记忆等约束的影响,这些约束限制了状态信息对行为引导的程度,从而限定了策略复杂度。标准的强化学习模型通常在不显式表示这些内部成本的情况下优化奖励,因此作为生物智能模型时适用性较低。我们推导出MI-SARSA,一种通过学习的边际动作先验和对该先验的状态特定偏差惩罚来引入互信息正则化的同策略时序差分算法。这产生了一个序列学习模型,其中状态信息在其预期回报收益足以证明额外信息成本合理时被选择性使用。关键的是,控制策略压缩的同一状态特定信息成本也生成了反应时间的试次水平预测,这使MI-SARSA区别于大多数预测选择或回报而非潜伏期的强化学习模型。在实证中,MI-SARSA产生了奖励-复杂度权衡,且更强的信息惩罚会产生更简单的策略,具有更低的控制成本和更快的反应时间。在环境变化下,增加正则化会减少切换后的性能下降,但也降低了渐近回报,揭示了鲁棒性-容量权衡。这些结果共同将MI-SARSA定位为认知约束下有界序列学习的模型。
英文摘要
Biological agents do not learn under conditions of unlimited computation. For humans, learning and choice are shaped by constraints on perception, attention, and working memory, which limit how much state information guides behavior and therefore bound policy complexity. Standard reinforcement learning models typically optimize reward without explicitly representing these internal costs, making them less suitable as models of biological intelligence. We derive MI-SARSA, an on-policy temporal-difference algorithm that incorporates mutual-information regularization through a learned marginal action prior and a penalty on state-specific deviations from that prior. This yields a sequential learning model in which state information is used selectively when its expected return benefit justifies the added informational cost. Critically, the same state-specific information cost that governs policy compression also generates trial-level predictions for reaction time, distinguishing MI-SARSA from most reinforcement learning models, which predict choices or returns but not latency. Empirically, MI-SARSA produces a reward-complexity tradeoff, and stronger information penalties produce simpler policies with lower control costs and faster reaction times. Under environment shift, increasing regularization reduces post-switch performance degradation but also lowers asymptotic return, revealing a robustness-capacity tradeoff. Together, these results position MI-SARSA as a model of bounded sequential learning under cognitive constraints.