发表机构
Microsoft; IIT Bhubaneswar(微软; 布巴内斯瓦尔印度理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对现有旅行规划基准未考虑不确定性的局限,推出UTP-Bench基准并设计三项评估指标,实验发现GPT-5等先进LLMs生成的旅行行程与人工行程存在显著差距。
AI 中文摘要
大型语言模型(LLMs)近期在自动生成旅行行程方面展现出强大能力。然而,现实世界的旅行规划本质上存在不确定性:交通延误、人流波动以及突发的随机延误常常会使原本可行的行程失效。现有基准如TravelPlanner和TripCraft均假设环境是确定性的,仅评估静态约束满足情况,而忽略生成的计划在面对此类不确定性时是否仍具备鲁棒性。为解决这一局限,我们推出UTP-Bench¹,这是一个面向不确定性感知旅行规划的大规模基准。该数据集整合了印度504个城市的真实旅行数据,涵盖景点、餐厅、住宿以及多式联运交通网络。为模拟现实中的干扰,UTP-Bench纳入了从主要城市收集的经验延误分布和人流密度模式,支持在随机条件下对旅行计划进行评估。我们进一步提出三项评估指标:缓冲区充足度评分(BAS)、人流感知时序评分(CATS)以及交通延误吸收评分(TDAS),用于量化生成的行程在应对交通延误和人流变化时保持鲁棒性的能力。对GPT-5、Qwen3、Mistral和Phi-4等先进LLMs开展的实验显示,模型生成的行程与人工编写的行程之间存在显著差距,尤其体现在时间缓冲、感知延误的交通调度以及人流敏感规划方面。
英文摘要
Large Language Models (LLMs) have recently demonstrated strong capabilities in automated travel itinerary generation. However, real- world travel planning is inherently uncertain: transportation delays, crowd fluctuations, and unexpected stochastic delays frequently inval- idate otherwise feasible schedules. Existing benchmarks like TravelPlanner and TripCraft assume deterministic environments, evaluating only static constraint satisfaction and ignoring whether generated plans remain robust when such uncertainties arise. To address this limitation, we introduce UTP-Bench1 , a large-scale benchmark for uncertainty-aware travel planning. The dataset integrates real-world travel data spanning 504 cities of India, including attractions, restau- rants, accommodations, and multi-modal trans- portation networks. To model realistic disrup- tions, UTP-Bench incorporates empirical delay distributions and crowd-density patterns col- lected from major cities, enabling evaluation of travel plans under stochastic conditions. We further propose three evaluation metrics, namely Buffer Adequacy Score (BAS), Crowd- Aware Timing Score (CATS), and Transport Delay Absorption Score (TDAS), which quan- tify the ability of generated itineraries to main- tain robustness against transit delays and crowd variability. Experiments with state-of-the-art LLMs like GPT-5, Qwen3, Mistral and Phi-4 re- veal substantial gaps between model-generated and human-authored plans, particularly in tem- poral buffering, delay-aware transportation scheduling, and crowd-sensitive planning.
Comments34 pages, 12 figures, 16 Tables, EMNLP 2026