arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

STRIDE:跨多样情境的文本到轨迹对齐自动化评估

STRIDE: Automated Evaluation of Text-to-Trajectory Alignment across Diverse Contexts

Wanchun Ni, Tao Qi, Leonel Aguilar, Jiugeng Sun, Marlene Wagner, Verena Zimmermann, Mennatallah El-Assady

arXiv 2609.34799首次发表:更新:

发表机构

ETH Zurich; Beijing University of Posts and Telecommunications(苏黎世联邦理工学院; 北京邮电大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

STRIDE提出首个自动化评估框架,通过VRDST协议、行为问题分解和确定性测量库,实现无需人类数据的文本到轨迹情境对齐评估,并构建STRIDE-Bench基准,验证其与人类判断80%的一致性。

AI 中文摘要

语言条件化的轨迹生成已经出现,但其评估方法却未能跟上步伐。现有的行人轨迹度量标准将轨迹与真实世界的人类数据进行比较。这无法扩展到跨多样情境的文本到轨迹生成,因为为每个场景收集人类轨迹成本高昂且不可行。此外,行人行为具有异质性和情境依赖性,没有单一度量标准能作为正确答案,而当前的评估框架也无法迁移到这一领域。这些挑战使得可扩展、可靠的评估变得困难。我们引入了STRIDE,这是首个用于评估场景描述与行人轨迹之间情境对齐的框架。STRIDE通过三个设计选择应对这些挑战。首先,我们从社会学理论中推导出我们的VRDST评估协议,以定义一个完整的评估空间。其次,它将高层情境分解为适应场景的行为问题。第三,每个问题都通过一个确定性的测量工具库来解决,该库能产生可复现的答案。综上,STRIDE能够在不需要人类轨迹数据的情况下,实现跨多样情境的完整、可验证、自动化且可扩展的评估。我们在人群领域将STRIDE实例化为STRIDE-Bench,包含1K个场景、6K个行为问题和11K个测量,并在30个真实世界地图上配有校准的预期答案。全面的人类验证表明,STRIDE-Bench与人类行为和判断一致,达到了80%的人类一致性。我们进一步评估了多个文本到轨迹模型,发现其情境对齐能力有限,且在细粒度情境条件化方面存在持续挑战。我们相信,STRIDE框架为情境对齐的行人轨迹生成的原则性评估迈出了第一步。

英文摘要

Language-conditioned trajectory generation is here, but its evaluation has not kept pace. Existing pedestrian trajectory metrics compare trajectories with real-world human data. This does not scale to text-to-trajectory generation across diverse contexts, as collecting human trajectories for every scenario is costly and infeasible. Moreover, pedestrian behavior is heterogeneous and context-dependent, with no single metric as the correct answer, and current evaluation frameworks are not transferable to this domain. These challenges make scalable, reliable evaluation difficult. We introduce STRIDE, the first framework for evaluating context alignment between scenario descriptions and pedestrian trajectories. STRIDE addresses these challenges through three design choices. First, we derive our VRDST evaluation protocol from sociological theories to define a complete evaluation space. Second, it decomposes high-level context into scenario-adaptive behavioral questions. Third, every question is resolved against a deterministic measurement tool library that yields reproducible answers. Together, STRIDE enables complete, verifiable, automated, and scalable evaluation across diverse contexts without requiring human trajectory data. We instantiate STRIDE in the crowd domain as STRIDE-Bench, comprising 1K scenarios, 6K behavioral questions, and 11K measurements with calibrated expected answers across 30 real-world maps. Comprehensive human validations show that STRIDE-Bench is consistent with human behavior and judgment, achieving 80% human agreement. We further evaluate several text-to-trajectory models, finding limited context-alignment capability and persistent challenges in fine-grained context conditioning. We believe that the STRIDE framework provides a first step toward principled evaluation of context-aligned pedestrian trajectory generation.

CommentsAccepted at NeurIPS 2026, Evaluations & Datasets Track

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑