LiDAR语言模型真的理解时空关系吗?
Do LiDAR Language Models Really Understand Spatio-temporal Relationships?
- Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院)
- Hunan University(湖南大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出LiDAR-Hallu基准,通过配对问题和关系特定分析揭示4D LiDAR语言模型在时空推理中依赖候选时长等捷径,无法真正区分物理关系。
AI中文摘要:
近期的4D LiDAR语言模型旨在推理对象及其不断演化的空间关系。然而,在我们的评估中,始终选择同一选项的准确率几乎与两个基于B4DL的配置的多项选择准确率相当。我们引入了LiDAR-Hallu,一个几何参考基准和诊断协议,包含150个nuScenes场景中的10,000个问题。它涵盖了对象存在性、自我相对位置、距离排序、相对运动和时序定位,并明确规定了选择对象、比较时间和确定参考答案的规则。我们的协议结合了固定答案和候选内容控制、具有相同提示但参考答案相反的跨场景配对,以及关系特定召回率。对100,000条记录响应的分析揭示了被总体准确率掩盖的失败。仅候选持续时间就使得时序答案无需观察LiDAR即可预测。在配对问题上,模型对需要相反答案的场景经常给出相同答案。关系特定分析进一步表明,两种配置在所有测试条件下都遗漏了每一个正横向运动案例。时序洗牌对比解码几乎没有带来净改进,因为修复在很大程度上被新错误所抵消,主要失败仍然存在。这些结果表明,评估时空推理需要测试模型是否区分所查询的物理关系,而不是仅仅依赖单个答案的准确率。源代码、检查点和数据已在此https URL发布。
英文摘要:
Recent 4D LiDAR language models aim to reason about objects and their evolving spatial relationships. Yet, in our evaluation, always selecting the same option nearly matches the multiple-choice accuracy of two B4DL-derived configurations. We introduce LiDAR-Hallu, a geometry-referenced benchmark and diagnostic protocol with 10,000 questions across 150 nuScenes scenes. It covers object existence, ego-relative position, distance ordering, relative motion, and temporal localization, with explicit rules for selecting objects, comparing times, and determining reference answers. Our protocol combines fixed-answer and candidate-content controls, cross-scene pairs with identical prompts but opposite reference answers, and relation-specific recall. Analysis of 100,000 recorded responses reveals failures hidden by aggregate accuracy. Candidate duration alone makes temporal answers predictable without observing LiDAR. On paired questions, the models frequently give the same answer to scenes requiring opposite answers. Relation-specific analysis further shows that both configurations miss every positive lateral-motion case across all tested conditions. Temporal-shuffle contrastive decoding provides little net improvement, as repairs are largely offset by new errors and the main failures persist. These results show that evaluating spatio-temporal reasoning requires testing whether models distinguish the queried physical relationships, rather than relying on individual-answer accuracy alone. The source code, checkpoints, and data are released at https://github.com/Awesome4D/4DMLLM_Hallucination_Bench.