arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

顺其自然:序贯决策中的等待学习

Let it Cook: Learning to Wait in Sequential Decision Making

Christopher Watson, Arjun Krishna, Dinesh Jayaraman, Rajeev Alur

arXiv 2608.11511首次发表:更新:

AI 中文总结

该研究针对序贯决策中智能体频繁感知决策的问题,提出采用带字典序目标的强化学习训练等待策略,在7类任务上实现了可观的等待时长,且不牺牲任务性能。

AI 中文摘要

在序贯决策中,智能体通常会在每个时间步观察环境并采取行动。然而,这种主动参与并非总是必要的;诸如煮咖啡之类的任务包含一些阶段,在这些阶段中,让环境自然演变而无需持续监控和控制也能达到同等效果。在这些阶段,智能体可以简单地等待以节省资源,或将注意力转移到另一项任务上。我们利用这些机会,通过训练“等待策略”来决定在何处以及等待多长时间。这包括放弃感知以执行等待动作,即设定一定时间步的刻意暂停。我们将“学习等待”形式化为在不牺牲任务性能(例如完成任务的总时间)的前提下,最小化感知和决策的频率。为训练等待策略,我们提出了一种采用字典序排序目标的强化学习方法。在4个离散状态的家庭任务和3个连续状态环境的实验中,我们表明我们的方法成功学习到了等待行为,并且可以调整预训练策略以在适当的地方等待。虽然不同任务允许在不牺牲任务性能的情况下进行不同时长的等待,但我们的方法始终能找到具有显著等待时间的解决方案,有时等待时间超过任务总时长的50%。

英文摘要

In sequential decision making, an agent typically observes its environment and acts at every timestep. However, such active participation may not always be necessary; tasks such as brewing coffee include periods that are served equally well by letting the environment evolve without constant monitoring and control. During such periods, the agent could simply wait to conserve its resources, or redirect its attention to another task. We capitalize on these opportunities by training a "waiting policy" that decides where and how long to wait. This involves forgoing sensing to commit to a wait action, representing a deliberate pause for a set number of timesteps. We formalize "learning to wait" as minimizing the frequency of sensing and decision making without sacrificing task performance (e.g., the total amount of time to complete a task). To train a waiting policy, we propose an approach that employs reinforcement learning with lexicographically ordered objectives. In experiments across 4 discrete-state household tasks and 3 continuous-state environments, we show that our approach successfully learns waiting behaviors, and can adapt pre-trained policies to wait where appropriate. While different tasks permit different amounts of waiting without sacrificing task performance, our approach consistently finds solutions with significant waiting, sometimes waiting for over 50 percent of the task duration.

Comments25 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑