arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TREK:面向复杂行程规划的大语言模型智能体旅行推理与评估工具包

TREK: A Travel Reasoning and Evaluation Kit for LLM Agents in Complex Trip Planning

Jinhu Qi, Wentao Zhang, Siu Man Ng, Feiyang Xu, Yanyu Chen, Yaoman Li, Irwin King

arXiv 2607.26977首次发表:更新:

发表机构

The Chinese University of Hong Kong; Macao Polytechnic University(香港中文大学; 澳门理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究推出TREK旅行基准测试,解决现有行程规划智能体基准的不足,实验显示最强LLM智能体仅在不足五成任务生成可行计划,满足旅行者未明确需求是普遍瓶颈。

AI 中文摘要

行程规划对使用工具的大语言模型(LLM)智能体而言是一项严苛的压力测试:可用的行程是单一产物,必须同时在多个维度上符合要求——每趟航班、酒店和景点必须存在且可预订,日期必须在物理上可通行,总费用必须符合预算,且计划需满足仅部分被明确的旅行者需求。现有的智能体基准测试分别对这些属性进行奖励,并通过软性规则或LLM评判标准对最终输出打分,无法验证返回的计划是否可执行,也不具备可复现性和可审计性。我们推出TREK(Travel Reasoning and Evaluation Kit),这是一个用于可行行程合成的基准测试:生成的单一计划需同时满足约束正确、无幻觉、时空可执行、预算有效且能响应旅行者未明确的角色需求。TREK包含800个多约束任务——其中533个可行,267个经证明不可行,且标注了路线/实体/预算方面的原因——基于一个由375个城市、13种角色构成的合成内部一致知识库,共212530条记录,通过经验证的RESTful API的生产级工具沙盒提供服务。每个任务由完全确定性的基于规则的评估器打分,无LLM评判,且附带经人工验证的黄金参考,该参考在同一评估器下得分为1.0,因此明确存在可达到的上限,剩余差距均为智能体的局限而非评分者的严格要求。在9个约束维度上对15个LLM智能体进行评估后,我们发现即使是最强的智能体(GPT-5.6)也仅在46.2%的可解任务上生成完全可行的计划,中位数为6.6%,最低值为0.0%;满足旅行者未明确的需求是普遍存在的瓶颈,即使在前沿水平也未得到解决。我们发布了该数据集、工具沙盒、确定性评估器和智能体代码,作为一个完全可复现的基准测试。

英文摘要

Travel planning is a demanding stress test for tool-using LLM agents: a usable itinerary is a single artifact that must be right along many axes at once - every flight, hotel, and attraction must exist and be bookable, the days must be physically traversable, the total must clear a budget, and the plan must serve a traveler whose needs are only partly stated. Existing agent benchmarks reward these properties one at a time and grade the final output with soft or LLM-judged rubrics, which cannot certify that a returned plan is executable and are neither reproducible nor auditable. We introduce TREK (Travel Reasoning and Evaluation Kit), a benchmark for feasible itinerary synthesis: producing a single plan that is jointly constraint-correct, hallucination-free, spatio-temporally executable, budget-valid, and responsive to the traveler's unstated persona needs. TREK comprises 800 multi-constraint tasks - 533 feasible and 267 provably infeasible with typed route/entity/budget causes - over a synthetic, internally consistent knowledge base of 212,530 records across 375 cities and 13 personas, served through a production-style tool sandbox of validated RESTful APIs. Every task is scored by a fully deterministic, rule-based evaluator with no LLM judge and ships a human-verified gold reference that scores a perfect 1.0 under that same evaluator, so the ceiling is demonstrably achievable and every remaining gap is an agent limitation rather than scorer strictness. Evaluating 15 LLM agents across nine constraint dimensions, we find that even the strongest (GPT-5.6) produces a fully-feasible plan on only 46.2% of solvable tasks, with a median of 6.6% and a floor of 0.0%; satisfying travelers' unstated needs emerges as the universal bottleneck, unsolved even at the frontier. We release the dataset, tool sandbox, deterministic evaluator, and agent code as a fully reproducible benchmark.

CommentsCode, data, and evaluator: https://github.com/TonyQJH/TREK-A-Travel-Reasoning-and-Evaluation-Kit-for-LLM-Agents-in-Complex-Trip-Planning

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

相关深度报道

↑