发表机构
School of Computer Science, Carnegie Mellon University; Lavoro AI; Stack AV; Bosch Center for Artificial Intelligence(卡内基梅隆大学计算机科学学院; 拉沃罗人工智能公司; 斯塔克自动驾驶公司; 博世人工智能中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出ControlledShifts框架与基准套件,通过标准化划分轨迹数据集及统一鲁棒性评分,评估不同容量Transformer架构在分布偏移下轨迹预测的鲁棒性,揭示其处理潜在相关性与环境结构的差异。
AI 中文摘要
轨迹预测是自动驾驶安全的核心,但基于学习的预测器在遇到训练数据未充分覆盖的场景时往往会急剧性能下降。许多方法试图通过以数据为中心或测试时自适应方法缓解分布偏移导致的性能下降,但这些方法通常仅在零散的泛化轴上进行验证,使得该领域缺乏标准化的方式来比较模型在可能遇到的各类偏移下的鲁棒性。为解决这一问题,我们提出了ControlledShifts,这是一个框架和基准套件,通过共享的表征与划分公式,系统地将现有轨迹数据集重新划分为分布内(已见)和分布外(未见)划分:其中表征函数确定基准所探测的变化轴,划分函数确定该轴尾部数据的保留方式。该套件包含三个针对关键拓扑和行为分布偏移的基准。此外,为聚合这些基准的多维性能指标,我们提出了一个统一鲁棒性评分,该评分从两个互补维度评估模型:预测质量(相对性能增益)和预测稳定性(偏移下的性能保持)。我们通过对知名的基于Transformer的架构进行基准测试来展示ControlledShifts,揭示了不同容量模型在处理潜在相关性和环境结构方面的关键差异。
英文摘要
Trajectory prediction is central to safety in autonomous driving, yet learning-based predictors tend to degrade sharply when encountering scenarios poorly represented by their training data. Many methods attempt to mitigate distribution shift degradation through data-centric or test-time adaptation approaches; however, they are typically validated along fragmented axes of generalization, leaving the field without a standardized way to compare robustness across shifts a model may encounter. To address this, we introduce ControlledShifts, a framework and benchmark suite that systematically re-splits existing trajectory datasets into in-distribution (seen) and out-of-distribution (unseen) partitions, via a shared characterization-and-splitting formulation, in which a characterization function fixes the axis of variation a benchmark probes and a splitting function fixes how the tail of that axis is withheld. The suite comprises three benchmarks targeting key topological and behavioral distribution shifts. Furthermore, to aggregate multi-dimensional performance metrics across these benchmarks, we propose a unified robustness score that evaluates models along two complementary dimensions: prediction quality (relative performance gain) and prediction stability (performance preservation under shift). We showcase ControlledShifts by benchmarking prominent transformer-based architectures, exposing critical differences in how models of varying capacities handle latent relevance and environmental structure.
Comments8 pages, 8 figures, 1 table