arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.18672cs.ROcs.AI

带不确定时变奖励的定向越野问题:面向日常服务机器人的框架与基准

Orienteering Problem with Uncertain Time-Varying Rewards: Framework and Benchmark for Everyday Service Robotics

Masafumi Endo, Kohei Honda, Yuu Jinnai, Ryo Yonetani

首次发表
浏览论文内容

中文总结 AI 辅助

针对日常服务机器人的带不确定时变奖励的定向越野问题,本文提出相关框架与基准,采用三种不同规划器并验证了长时域在线适应规划的有效性。

中文摘要 AI 辅助

本文提出了带不确定时变奖励的定向越野问题(OP-UTVR),这是定向越野问题(OP)的一种新型变体。现有大多数OP公式均假设奖励为预先已知,而实际应用中奖励是不确定且随时间变化的,例如配送机器人面临的客户需求波动。OP-UTVR放宽了这一假设,允许智能体通过观测估计奖励动态并预测未来奖励,从而在奖励随机变化和不可避免的预测误差下做出知情的路径规划决策。我们采用三种规划器解决该问题,它们在规划时域和在线适应性方面存在差异,并推导了这些规划器在奖励随机性下的性能理论界。我们还为OP-UTVR引入了移动服务机器人基准,机器人需在室内环境中穿行于行人之间。实验揭示了规划时域与适应性之间的权衡关系,并证明了带在线适应性的长时域规划的有效性。

英文摘要

We present the orienteering problem with uncertain time-varying rewards (OP-UTVR), a novel variant of the orienteering problem (OP). While most existing OP formulations assume rewards to be known in advance, practical applications involve uncertain and time-varying rewards, as with shifting customer demand for delivery agents. OP-UTVR relaxes this assumption by allowing agents to estimate reward dynamics from observations and forecast future rewards. This enables informed routing decisions despite stochastic reward changes and inevitable prediction errors. We address this problem using three planners that differ in planning horizon and online adaptivity, and derive theoretical bounds on their performance under reward stochasticity. We further introduce a mobile service robot benchmark for OP-UTVR, where a robot navigates among pedestrians in indoor environments. Experiments reveal trade-offs between planning horizon and adaptivity, and demonstrate the effectiveness of long-horizon planning with online adaptation.

补充信息

↑