arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.34440cs.CVcs.RO

当分数成为目标:重新思考自动驾驶中的指标有效性

When the Score Becomes the Target: Rethinking Metric Validity in Autonomous Driving

  • University of North Texas(北德克萨斯大学)
  • Toyota InfoTech Labs(丰田信息技术实验室)

机构由 AI 辅助整理,请以论文原文为准。

Morui Zhu, Deyuan Qu, Qi Chen, Kentaro Oguchi, Qing Yang

AI总结:

本研究探讨自动驾驶基准分数被优化后是否仍能可靠反映驾驶改进,通过分解评分过程并实验发现分数提升可能不持久,强调需同时关注测量内容与执行方式。

AI中文摘要:

驾驶基准分数越来越多地被用于评估,同时也作为优化目标。这引发了一个根本性问题:一旦分数本身被优化,分数提升是否仍然是驾驶改进的可靠证据?我们通过考察评分过程如何响应驾驶行为的变化,以及由此产生的提升在重复执行和重新规划下是否持续,来解决这个问题。我们将该过程分解为执行、测量、子分数映射和聚合。受控干预揭示了显著的行为变化,但由于请求运动与执行运动之间的差异被省略、阈值化或衰减,这些变化得到的分数响应很小。闭环比较进一步表明,当执行接口改变时,优化收益可能逆转,这证明了它们对请求如何被执行并作为反馈返回的依赖性。总之,这些发现将指标保留的行为差异与其收益转移的条件联系起来。因此,优化下的指标有效性需要同时考察评分过程测量什么以及优化行为如何被执行。

英文摘要:

Driving benchmark scores are increasingly used not only for evaluation but also as optimization targets. This raises a fundamental question: do score gains remain reliable evidence of driving improvement once the score itself is optimized? We address this question by examining how the scoring process responds to changes in driving behavior and whether the resulting gains persist under repeated execution and replanning. We decompose the process into execution, measurement, subscore mapping, and aggregation. Controlled interventions reveal substantial behavioral changes that receive little score response because distinctions are omitted, thresholded, or attenuated between requested and executed motion. Closed-loop comparisons further show that optimization gains can reverse when the execution interface changes, demonstrating their dependence on how requests are executed and returned as feedback. Together, these findings connect the behavioral distinctions preserved by a metric to the conditions under which its gains transfer. Metric validity under optimization therefore requires examining both what the scoring process measures and how the optimized behavior is executed.

补充信息

↑