RideSkill:一种基于大语言模型驱动自动演化的通用拼车分层算法
RideSkill: A Hierarchical Algorithm for Generalized Ride Sharing with LLM-Driven Automatic Evolution
浏览论文内容
中文总结 AI 辅助
针对现有拼车算法泛化性差、推理需频繁调用LLM的问题,本文提出RideSkill分层算法,通过LLM自动演化训练组件,实现支持车辆共享的高效拼车调度,保障实时性能。
中文摘要 AI 辅助
拼车允许多个具有不同起讫点(OD)的乘客共享同一车辆,是一个具有挑战性的运营问题,需要在不确定且多变的场景下,高效地将不同OD的订单捆绑并分配给车辆。尽管多智能体强化学习(MARL)方案已取得良好性能,但存在泛化能力有限(难以适配不同环境场景)、迁移性低(难以适配不同平台目标)以及大规模系统训练困难(如维度灾难)等问题。近期,受大语言模型(LLM)规模扩展的启发,多项研究将LLM应用于叫车系统,要么直接将LLM作为决策智能体,要么用于自动算法设计。然而,这些方法均不支持车辆共享,该问题会使状态和动作空间呈指数级扩展,增加复杂度;且多数方法在推理时需要频繁调用LLM,无法实现实时部署。为解决上述问题,本文提出RideSkill,一种利用LLM辅助自动算法设计的拼车分层方法。RideSkill包含两个核心组件:一是组合器,从学习到的技能库中为每辆车分配合适的技能,以实现不同场景和目标下的自适应调度;二是重定位器,将闲置车辆依次重新部署到新兴区域,避免车辆间冲突。关键在于,技能库、组合器和重定位器均通过基于LLM的自动演化方法训练,部署期间无需调用LLM,确保了高实时性能。
英文摘要
Ride-sharing, which allows multiple passengers with different origin-destination (OD) pairs to share a single vehicle, is a challenging operational problem, as it requires orders with different OD pairs to be efficiently bundled and assigned to vehicles under uncertain and varying scenarios. Although multi-agent reinforcement learning (MARL) solutions have achieved promising performance, they suffer from limited generalization (adapting to different environmental scenarios), low transferability (adapting to different platform objectives), and training difficulties in large-scale systems, such as the curse of dimensionality. Recently, motivated by the scaling of large language models (LLMs), several works have incorporated LLMs into ride-hailing systems, either by employing LLMs directly as decision-making agents or using them for automatic algorithm design. However, none of these approaches support vehicle sharing, which complicates the problem by expanding both the state and action spaces exponentially. Moreover, most of them require frequent LLM calls at inference time, making them infeasible for real-time deployment. To address these issues, we propose RideSkill, a hierarchical method for ride-sharing that leverages LLM-assisted automatic algorithmic design. RideSkill consists of a combiner that assigns appropriate skills to each vehicle from a learned skill repository, enabling adaptive dispatch under varying scenarios and objectives, and a repositioner that sequentially relocates idle vehicles to emerging regions, avoiding conflicts among vehicles. Crucially, the skill repository, combiner, and repositioner are all trained by an LLM-based automatic evolutionary method, eliminating the need for LLM calls during deployment and thus ensuring high real-time performance.
发表机构
- The Hong Kong University of Science and Technology(香港科技大学)
- Noah’s Ark Lab, Huawei(华为诺亚方舟实验室)
- The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
机构由 AI 辅助整理,请以论文原文为准。