arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

不确定性下延迟奖励获取的最优性

Optimality of delayed reward attainment under uncertainty

Chuwei Wang, Alexander Vladimirsky, Anastasia Bizyaeva

arXiv 2610.03506首次发表:更新:

发表机构

Cornell University(康奈尔大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文研究不确定性下延迟奖励获取的最优性,提出确定性最优控制框架,证明延迟决策在难区分备选方案且等待成本低时最优,并给出缩短规划范围的条件及策略性感知移动策略。

AI 中文摘要

在不确定性下进行决策需要在额外信息的价值与延迟决策至信息可用之间的成本之间进行权衡。这一挑战在具身决策中尤为普遍,例如动物在食物源之间移动以及机器人选择导航目标时所面临的决策,在这些决策中,通过空间的移动既影响信息的质量,也影响行动的成本。我们考虑该问题的一个版本,其中决策是在不确定价值的奖励之间进行选择,而这些奖励只能通过充分改变决策者的状态来获得。我们的重点是奖励学习与状态演化之间的相互作用,并将其重述为一个具有两个终端备选方案的确定性最优控制问题,其中一个备选方案具有已知奖励,另一个则通过噪声观测逐渐学习。我们证明,在许多系统中,延迟奖励获取通常是最优的,尤其是当备选方案最初难以区分且等待成本相对较低时。对于一部分问题,我们还推导了充分条件,使得可以先验地缩短规划范围,保证进一步的观测不会影响价值函数。最后,当观测质量依赖于状态时,我们表明,控制器在做出决策之前,战略性地将状态演化到信息更丰富的感知位置通常是最优的,这与在生物和工程系统中广泛观察到的行为一致。

英文摘要

Decision making under uncertainty requires balancing the value of additional information against the cost of delaying the decision until that information becomes available. This challenge is especially prevalent for embodied decisions such as those faced by animals moving between food sources and robots selecting navigation targets, where movement through space shapes both the quality of information and the cost of acting. We consider a version of this problem, where the decision is on choosing among rewards of uncertain value, and those rewards can only be attained through changing the decision maker's state sufficiently. Our focus is on interactions between reward learning and state evolution, restated as a deterministic optimal control problem with two terminal alternatives, one with a known reward and the other learned gradually through noisy observations. We demonstrate that delayed reward attainment is often optimal in a variety of systems, especially so when the alternatives are initially difficult to distinguish and the cost of waiting is relatively low. For a subset of problems, we also derive sufficient conditions that allow shortening the planning horizon a priori, guaranteeing that further observations would not affect the value function. Finally, when observation quality depends on the state, we show that it is often optimal for a controller to evolve the state strategically toward more informative sensing locations before making a decision, in line with the behavior broadly observed in biological and engineered systems.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑