arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

分布偏移下异构多机器人任务分配的模型强化学习方法

Model-Based Reinforcement Learning for Heterogeneous Multi-Robot Task Assignment Under Distribution Shifts

Daniel Garces, Sara Castro, Adrian Haimovich, Byron Crowe, Stephanie Gil

arXiv 2608.21554首次发表:更新:

发表机构

John A. Paulson School Of Engineering And Applied Sciences, Harvard University; Harvard Medical School; Beth Israel Deaconess Medical Center; Stanford University School of Medicine(哈佛大学约翰·A·保尔森工程与应用科学学院; 哈佛医学院; 贝斯以色列女执事医疗中心; 斯坦福大学医学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出一种预测感知自适应滚动框架,用于分布偏移下异构多机器人任务分配,在医院护理任务案例中,该方法较多种基线缩短了等待时间,提升了服务覆盖与尾部延迟性能。

AI 中文摘要

异构多机器人服务系统需将请求分配给兼容机器人、构建可行调度方案,并在新任务在线到达时进行自适应调整。历史数据可帮助预测未来需求,但过度依赖不准确的预测会在分布偏移下降低性能。针对包含计划内和实时请求的异构多机器人任务分配问题,本文提出一种预测感知自适应滚动框架。该问题被建模为有限时间随机动态规划,融合了机器人-任务兼容性、有序服务要求、路由约束、服务窗口及时间终点回报要求。所提策略通过采样未来请求场景评估当前分配,同时限制仅对已观测到的请求做出即时承诺。为支持在线使用,该框架结合剪枝候选控制、等待动作及交互感知基础策略以实现高效的未来成本估计。通过基于近期预测不匹配自适应重加权预测请求,并选择性重新优化已分配但未启动的请求,提升了对预测误差的鲁棒性。本文还提出一种历史数据驱动的方法,用于部署前选择异构机器人编队组成。在使用医院住院楼层真实护理任务请求的案例研究中,与反应式、令牌传递、预测定位及近视贪心基线方法相比,所提方法实现了接近完全的服务覆盖,并缩短了已服务请求的等待时间,在尾部延迟指标上取得了最大改进。

英文摘要

Heterogeneous multi-robot service systems must assign requests to compatible robots, construct feasible schedules, and adapt as new tasks arrive online. Historical data can help anticipate future demand, but relying too heavily on inaccurate predictions can degrade performance under distribution shifts. We develop a prediction-aware adaptive rollout framework for heterogeneous multi-robot task assignment with scheduled and real-time requests. The problem is formulated as a finite-horizon stochastic dynamic program incorporating robot-task compatibility, ordered service requirements, routing constraints, service windows, and end-of-horizon return requirements. The proposed policy evaluates current assignments using sampled future request scenarios while restricting immediate commitments to requests already observed. To enable online use, the framework combines pruned candidate controls, wait actions, and an interaction-aware base policy for efficient future-cost estimation. Robustness to forecast error is provided by adaptively reweighting predicted requests based on recent prediction mismatch and selectively re-optimizing assigned but unstarted requests. We also introduce a historical-data-driven procedure for selecting the heterogeneous fleet composition before deployment. In a case study using real nursing-task requests from hospital inpatient floors, the proposed approach achieves near-complete service and reduces serviced-request wait times relative to reactive, token-passing, prediction-positioning, and myopic greedy baselines, with the largest improvements in tail-delay metrics.

Comments34 pages, 14 figures, 4 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑