RISE:跨三维跟踪与结构化视觉-语言推理的路边基础设施序列理解
RISE: Roadside Infrastructure Sequence Understanding across 3D Tracking and Structured Vision-Language Reasoning
浏览论文内容
中文总结 AI 辅助
本研究提出RISE框架,结合仅图像的三维跟踪方法与结构化视觉-语言推理,构建含33910个问答对的RISE-VQA数据集,通过RISE-Bench评估任务,揭示相关挑战并验证域自适应与时间上下文的益处。
中文摘要 AI 辅助
我们提出了RISE(Roadside Infrastructure Sequence Understanding and Evaluation,路边基础设施序列理解与评估),这一框架涵盖路边序列中的度量三维跟踪与结构化视觉-语言推理。对于度量跟踪,我们仅基于图像的方法将SAM3视频身份与校准引导的掩码一致性相结合,用于多视图身份关联,无需激光雷达(LiDAR)或特定任务的三维训练即可恢复持久的三维跟踪。其校准条件下的几何特性使该过程可在不同校准的多摄像头路口实例化,无需针对特定布局进行重新训练。在来自6个路口的20个人工审核的片段上,生成的跟踪在定义的多视图评估范围内达到了66.9的MOTA。对于结构化视觉-语言推理,一个人工审核的MLLM流程挖掘高价值片段,并使用受约束的全上下文Oracle构建基于边界框(bbox)的预测问答(QA),同时不向被评估模型暴露未来证据。由此产生的RISE-VQA数据集包含来自16个路口、61个路边视角的557个片段中的33910个问答对。其路口保留的RISE-Bench使用确定性的特定任务指标评估语义选择、坐标、未来边界框和交互集。实验表明,领域自适应和一般情况下的时间上下文带来了持续的益处,同时也揭示了空间定位、未来定位和交互推理方面的持续挑战。
英文摘要
We present RISE (Roadside Infrastructure Sequence Understanding and Evaluation), a framework spanning metric 3D tracking and structured vision-language reasoning in roadside sequences. For metric tracking, our image-only method combines SAM3 video identities with calibration-guided mask agreement for multi-view identity association, recovering persistent 3D tracks without LiDAR or task-specific 3D training. Its calibration-conditioned geometry allows the procedure to be instantiated at different calibrated multi-camera intersections without layout-specific retraining. On 20 human-reviewed clips from six intersections, the generated tracks achieve 66.9 MOTA within the defined multi-view evaluation scope. For structured vision-language reasoning, a human-reviewed MLLM pipeline mines high-value clips and uses a constrained full-context Oracle to construct bbox-grounded predictive QA without exposing future evidence to evaluated models. The resulting RISE-VQA dataset contains 33,910 QA pairs from 557 clips across 16 intersections and 61 roadside views. Its intersection-held-out RISE-Bench evaluates semantic choices, coordinates, future boxes, and interaction sets with deterministic task-specific metrics. Experiments show consistent benefits from domain adaptation and generally from temporal context, while revealing persistent challenges in spatial grounding, future localization, and interaction reasoning.