arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ALIGN-HOLD:滴滴大规模网约车匹配中的实时保持控制的经验对齐

ALIGN-HOLD: Experience Alignment for Real-Time Hold Control in Large-Scale Ride-Hailing Matching at DiDi

Zuhao Zhang, Xu Liu, Kai Wan, Zihao Lu, Li Ma, Shuai Li

arXiv 2609.09685首次发表:更新:

发表机构

Shanghai Jiao Tong University; Didichuxing Co. Ltd(上海交通大学; 滴滴出行科技有限公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

ALIGN-HOLD通过从隐式市场偏好中学习保持策略,利用经验对齐框架和奖励模型,在滴滴平台上显著提升行程完成率和司机收入,减少乘客取消。

AI 中文摘要

实时保持控制是大型网约车系统中的高杠杆机制:通过选择性地延迟司机-订单配对,平台可以等待更好的匹配机会,从而提升端到端的乘客-司机体验。现有的生产系统如EXHOLD,从行程完成、取消、等待时间和司机付出的手工组合中学习基于bandit的保持策略。然而,随着市场偏好的异质性和观察到的乘客-司机行为可能稀疏、嘈杂且受动态供需条件影响,设计此类奖励变得越来越困难。我们提出了ALIGN-HOLD,一个生产规模的经验对齐框架,从隐式市场偏好中学习保持策略。ALIGN-HOLD从订单轨迹、司机轨迹和同期局部匹配图中构建互补的偏好对,并使用平衡多视图采样和模型自适应困难偏好采样训练经验奖励模型(RM)。在基于模拟器的策略学习过程中,冻结的RM提供密集的、上下文相关的奖励,并支持对低可辨识性交互进行无标签过滤,这些交互的行为反馈难以归因于匹配质量。我们在滴滴的网约车平台上部署了ALIGN-HOLD,并在为期28天的随机A/B实验中进行了评估,覆盖每天约100,000个乘客请求。与已部署的生产策略相比,ALIGN-HOLD在行程完成率和司机收入方面取得了统计显著的提升,同时在司机接受前后显著减少了乘客取消。互补的消融实验、RM诊断和行为分析验证了所提出组件的贡献。ALIGN-HOLD已全面上线,目前正在服务滴滴的巴西市场。

英文摘要

Real-time hold control is a high-leverage mechanism in large-scale ride-hailing systems: by selectively deferring driver-order pairs, the platform can wait for better matching opportunities and improve end-to-end passenger-driver experience. Existing production systems such as EXHOLD learn bandit-based hold policies from handcrafted combinations of trip completion, cancellations, waiting time, and driver effort. However, designing such rewards becomes increasingly difficult as marketplace preferences are heterogeneous and observed passenger-driver behavior can be sparse, noisy, and affected by dynamic supply-demand conditions. We present ALIGN-HOLD, a production-scale experience alignment framework that learns hold policy from implicit marketplace preferences. ALIGN-HOLD constructs complementary preference pairs from order trajectories, driver trajectories, and contemporaneous local matching graphs, and trains an experience Reward Model (RM) using balanced multi-view sampling and model-adaptive hard preference sampling. During simulator-based policy learning, the frozen RM provides a dense, context-dependent reward and supports label-free filtering of low-identifiability interactions whose behavioral feedback is difficult to attribute to matching quality. We deploy ALIGN-HOLD on DiDi's ride-hailing platform and evaluate it in a 28-day randomized A/B experiment, covering approximately 100,000 passenger requests per day. Compared with the deployed production policy, ALIGN-HOLD achieves statistically significant improvements in trip completion rate and driver income, while significantly reducing passenger cancellations before and after driver acceptance. Complementary ablations, RM diagnostics, and behavioral analyses validate the contributions of the proposed components. ALIGN-HOLD has been fully ramped up and is currently serving DiDi's Brazil marketplace.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑