arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.10897cs.LG

面向多平台调度优化的部分可观测学习

Partially Observable Learning for Multi-Platform Dispatch Optimization

Fengming Yao, Man Luo

首次发表
浏览论文内容

中文总结 AI 辅助

本文针对多平台即时配送调度的部分可观测问题,提出POLO框架,通过建模平台-网格智能体、注意力策略表示与反事实奖励塑形,在多场景下实现优于基线的调度性能

中文摘要 AI 辅助

即时配送平台已成为城市物流的关键组成部分,日益依赖众包骑手完成高度动态的订单。在实际系统中,骑手并非专属单个平台,可能同时服务多个平台,而受隐私和运营约束,各平台仅能观测自身订单与骑手交互,形成具有固有部分可观测性的多平台调度环境。然而,现有调度优化研究大多假设骑手完全可观测且必须接受分配,部署于实际多平台场景时会导致性能大幅下降。本文提出POLO,一种用于多平台即时配送系统调度优化的部分可观测多智能体强化学习框架。POLO首先将每个平台-网格对建模为独立智能体,仅从平台局部观测学习调度策略,使学习过程与实际隐私和运营约束一致;为支持在不完整且异构的骑手信息下的有效决策,POLO引入新型基于注意力的策略表示,选择性聚合骑手间信息;此外,设计反事实奖励塑形机制,缓解跨网格联合动作引发的非平稳性,实现更稳定、可扩展的学习。开发高保真模拟器,在不同平台数量和系统规模下评估调度性能,大量实验表明,POLO在平台收益和骑手出行效率上始终优于强基线,凸显其在实际多平台场景中的鲁棒性与有效性。

英文摘要

Instant delivery platforms have become a critical component of urban logistics, increasingly relying on crowdsourced couriers to fulfill highly dynamic orders. In real-world systems, couriers are not exclusive to a single platform and may concurrently serve multiple platforms, while each platform can only observe its own orders and couriers' interactions due to privacy and operational constraints. This results in a multi-platform dispatch environment with inherent partial observability. However, most existing works on dispatch optimization assume full courier observability and mandatory assignment acceptance, causing substantial performance degradation when deployed in realistic multi-platform settings. In this paper, we propose POLO, a partially observable multi-agent reinforcement learning framework for dispatching optimization in multi-platform instant delivery systems. POLO firstly models each platform-grid pair as an independent agent that learns dispatch policies solely from platform-local observations, aligning the learning process with real-world privacy and operational constraints. To support effective decision-making under incomplete and heterogeneous courier information, POLO introduces a novel attention-based policy representation that selectively aggregates inter-courier information. Moreover, we design a counterfactual reward shaping mechanism to mitigate the non-stationarity induced by joint actions across grids, leading to more stable and scalable learning. We develop a high-fidelity simulator to evaluate dispatch performance under varying numbers of platforms and system scales. Extensive experiments demonstrate that POLO consistently outperforms strong baselines in terms of platform revenue and courier travel efficiency, highlighting its robustness and effectiveness in realistic multi-platform settings.

发表机构

  • University of Exeter(埃克塞特大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑