arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.21323cs.CV

VeriFuse:面向协同3D感知的有界视觉-语言仲裁与推理引导精化

VeriFuse: Bounded Vision-Language Arbitration and Reason-Guided Refinement for Cooperative 3D Perception

  • Tsinghua University(清华大学)
  • MIT(麻省理工学院)
  • UESTC(电子科技大学)
  • Northeastern University(东北大学)
  • KAIST(韩国科学技术院)

机构由 AI 辅助整理,请以论文原文为准。

Hongyi Lin, Yiyao Liu, Qi Kang, Heye Huang, Yang Liu, Haris Koutsopoulos, Jinhua Zhao

AI总结:

VeriFuse提出有界仲裁框架,让冻结VLM在协同3D感知中仅做语义仲裁(选择/精化/拒绝),由确定性约束决定几何,在DAIR-V2X上实现0.494/0.357 AP50/AP70且延迟影响极小。

AI中文摘要:

视觉-语言模型(VLM)已在多种任务中展现出强大的场景理解和语义判断能力,但它们在协同感知中的恰当角色仍不明确。直接要求VLM回归3D检测结果既不可靠又计算昂贵,而使用它来选择单一来源的输出则会丢弃来自其他智能体的有用信息。我们提出了VeriFuse,一个用于车路协同3D检测的有界仲裁框架。每个智能体首先独立生成检测结果。围绕每个车辆和路侧提议,VeriFuse生成基于来源条件的几何候选,并将原始检测、其扰动以及跨来源假设合并到一个统一的候选池中。随后,一个冻结的VLM在三种可接受动作中进行选择:当存在合适的候选时选择(SELECT)该候选;当对象被支持但所有候选在几何上均不充分时,对现有锚点进行精化(REFINE);或拒绝(REJECT)一个不受支持的路侧专属提议。在DAIR-V2X数据集上的实验表明,VeriFuse实现了0.494/0.357的协同3D AP50/AP70,并将300毫秒延迟下的相对车辆侧BEV AP50下降限制在1.7%以内。总体而言,VeriFuse为VLM在协同感知中分配了一个清晰且受限的角色:语义推理解决跨智能体假设之间的歧义,而确定性约束决定最终的3D几何。

英文摘要:

Vision-language models (VLMs) have demonstrated strong scene understanding and semantic judgment across diverse tasks, but their appropriate role in cooperative perception remains unclear. Directly asking a VLM to regress 3D detections is unreliable and computationally expensive, whereas using it to select the output of a single source discards useful information from other agents. We introduce VeriFuse, a bounded arbitration framework for vehicle-infrastructure cooperative 3D detection. Each agent first produces detections independently. Around each vehicle and roadside proposal, VeriFuse generates source-conditioned geometric candidates and combines the original detections, their perturbations, and cross-source hypotheses into a unified candidate pool. A frozen VLM then chooses among three admissible actions: SELECT an adequate candidate; REFINE an existing anchor when an object is supported but all candidates are geometrically inadequate; or REJECT an unsupported infrastructure-only proposal. Experiments on the DAIR-V2X dataset show that VeriFuse achieves 0.494/0.357 cooperative 3D AP50/AP70 and limits the relative vehicle-side BEV AP50 drop under a 300 ms delay to 1.7%. Overall, VeriFuse assigns the VLM a clear and constrained role in cooperative perception: semantic reasoning resolves ambiguity among cross-agent hypotheses, while deterministic constraints determine the final 3D geometry.

补充信息

↑