发表机构
Tech LLC; University of Virginia; University of Massachusetts Amherst(529科技有限责任公司; 弗吉尼亚大学; 马萨诸塞大学阿默斯特分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对个性化设备端行程生成问题,提出PLA框架,分计划、学习、适应三阶段,通过构建规划器集合、拟合奖励模型及进行局部优化来生成行程,实验表明该框架在可行性和胜率上表现出色,还提升了生产部署中的行程完成率并降低延迟。
AI 中文摘要
生成个性化旅行行程是一项复杂的规划任务,在硬组合可行性和软潜在合意性之间存在矛盾。经典优化能执行约束但无法捕捉旅行者主观偏好,基于学习的方法能建模偏好却不能保证可行性,移动部署对两者都有额外资源限制。为此提出PLA框架,分三个阶段。计划阶段构建轻量级规划器的异构集合生成可行候选;学习阶段从行程比较中拟合紧凑奖励模型捕捉进度、地理连贯性等属性;适应阶段在设备感知计算预算内进行可行性保持的局部优化。在超100个美国城市的2519次成对人工比较中,奖励引导的集合实现67.8%胜率,比最佳单规划器高11.2个百分点且100%可行,前沿语言模型在相同约束下可行性为0%。奖励模型在留出城市上有67.6%平均留一城市准确率。在FlyEnJoy生产部署中,PLA将行程完成率提高91%,设备端平均延迟109.9毫秒。
英文摘要
Generating personalized trip itineraries is a complex planning task and involves a tension between hard combinatorial feasibility and soft latent desirability. Classical optimization enforces constraints but fails to capture subjective traveler preferences. While learning-based approaches model preferences, they cannot guarantee feasibility. Mobile deployment imposes additional resource constraints on both. To address this, we propose Plan, Learn, Adapt (PLA), a three-stage framework for personalized on-device itinerary generation. The Plan stage builds a heterogeneous ensemble of lightweight planners that produces structurally diverse feasible candidates. From pairwise itinerary comparisons, Learn fits a compact Bradley-Terry reward model that captures emergent schedule properties such as pacing, geographic coherence, and day balance, which per-POI signals miss. Finally, Adapt applies feasibility-preserving local refinement within a device-aware compute budget; every intermediate state is feasible by construction. On 2,519 pairwise human comparisons across more than 100 U.S. cities, the reward-guided ensemble achieves a 67.8% win rate, 11.2 percentage points above the best single planner, with 100% feasibility. Three frontier LLMs, GPT-5, Claude Opus 4.5, and Gemini 3 Pro, achieve 0% feasibility under the same constraints. The reward model generalizes across held-out cities, with a 67.6% mean leave-one-city-out accuracy. In production deployment within FlyEnJoy, PLA increased itinerary completion rates by 91%, with 109.9 ms average on-device latency.