重新规划、修复还是编辑?资源中断下旅行代理行程修订的统一实证评估
Replan, Repair, or Edit? A Unified Empirical Evaluation of Travel Agents for Itinerary Revision under Resource Disruptions
- Euler AI(欧拉人工智能公司)
- University of New South Wales(新南威尔士大学)
- University of Sydney(悉尼大学)
- Vecton AI(Vecton人工智能公司)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究系统比较了LLM-Z3重新规划、IPyHOPPER分层修复和iTIMO局部编辑三种行程修订方法,发现分层修复在保留已接受行程和计算效率上更优,为资源中断下的行程修订提供了实用指南。
AI中文摘要:
旅行规划代理生成的行程在用户接受后可能因航班取消、酒店不可用或景点关闭而变得不可行。修订这些行程涉及完全重新规划、经典计划修复和基于LLM的旅行代理修订,这些方法在任务表述和评估协议上的差异阻碍了比较。我们使用两个TREK衍生的基准集进行了系统性实证研究:500个单中断案例(包括可行和不可行实例)和200个可行的同时复合中断案例。我们比较了LLM-Z3完全重新规划、IPyHOPPER分层修复和iTIMO局部修订适配器在有效性、计划稳定性和计算成本方面的表现。使用Gemini的LLM-Z3在复合中断成功率上达到了观测到的最高值。IPyHOPPER在单中断整体成功率上几乎与该配置持平,同时在成功修复中保留了显著更多的已接受行程内容。成功的分层和局部修复比完全重新规划做出的编辑更少,并保留了更多已接受的承诺。计算特征有所不同:IPyHOPPER不使用LLM推理,评估的LLM-Z3适配器使用紧凑的单次调用推理,而iTIMO适配器消耗了显著更多的令牌。该研究为在评估设置中平衡可行性恢复、承诺保留和计算成本提供了实用指南。
英文摘要:
Travel-planning agents generate itineraries that may become infeasible after acceptance because of flight cancellations, hotel unavailability, or attraction closures. Revising these itineraries involves full replanning, classical plan repair, and LLM-based travel-agent revision, whose differing task formulations and evaluation protocols hinder comparison. We conduct a systematic empirical study using two TREK-derived benchmark sets: 500 single-disruption cases, including feasible and infeasible instances, and 200 feasible simultaneous compound-disruption cases. We compare LLM-Z3 full replanning, IPyHOPPER hierarchical repair, and an iTIMO local-revision adapter across effectiveness, plan stability, and computational cost. LLM-Z3 with Gemini achieved the highest observed compound-disruption success. IPyHOPPER nearly matched that configuration's single-disruption overall success, while preserving substantially more of the accepted itinerary on successful repairs. Successful hierarchical and local repairs made fewer edits and retained more accepted commitments than full replanning. Computational profiles differed: IPyHOPPER used no LLM inference, the evaluated LLM-Z3 adapter used compact one-call inference, and the iTIMO adapter consumed substantially more tokens. The study provides practical guidelines for balancing feasibility recovery, commitment preservation, and computational cost within evaluated settings.