发表机构
Korea Advanced Institute of Science and Technology(韩国科学技术院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
FloorSAV是一种注入动态二维平面图的新型框架,可提升AV-LLMs在SAVED-Bench等基准上的空间推理能力,为具身智能的空间推理提供了有效方案。
AI 中文摘要
在动态自我中心环境中的三维空间推理对于具身智能至关重要,但视听大语言模型(AV-LLMs)缺乏直接从原始感官流处理和内化全局几何信息的明确机制。现有方法要么需要昂贵的微调,要么未能充分利用模型的跨模态推理能力。本文提出FloorSAV,这是一种通过渲染动态二维平面图明确锚定空间视听上下文的新型框架。通过整合三维点云、相机轨迹、空间音频线索和语义锚定的对象地标,我们将该平面图作为与自我中心视频同步的流注入AV-LLM。AV-LLMs利用其多模态能力,在平面图解释指导下的单次推理中联合推理视觉、听觉和几何线索。我们进一步引入SAVED-Bench(带动态智能体的空间视听自我中心基准),构建现实场景中空间能力的核心任务:动态相对性、区域和路径推理问答。FloorSAV在SAVED-Bench和SAVVY-Bench的各类任务上提升了AV-LLMs的空间推理能力。基于真实平面图的研究证明,FloorSAV凭借准确的空间信息具有巨大潜力。
英文摘要
While 3D spatial reasoning in dynamic egocentric environments is crucial for embodied intelligence, audio-visual large language models (AV-LLMs) lack explicit mechanisms to process and internalize global geometry directly from raw sensory streams. Existing approaches either require costly fine-tuning or underutilize the model's cross-modal reasoning capacities. In this paper, we propose FloorSAV, a novel framework that explicitly grounds spatial audio-visual context by rendering a dynamic 2D floormap. By integrating 3D point clouds, camera trajectories, spatial audio cues, and semantically grounded object landmarks, we inject this floormap into the AV-LLM as a synchronized stream with an egocentric video. AV-LLMs utilize their multi-modal capabilities to jointly reason over visual, auditory, and geometric cues in a single inference with floormap interpretation guidance. We further introduce SAVED-Bench (Spatial Audio-Visual Egocentric Benchmark with Dynamic Agents), constructing essential tasks of spatial capability in real-world scenarios: dynamic relativity, regional, and path reasoning QAs. FloorSAV improves AV-LLMs' spatial reasoning on various tasks from both SAVED-Bench and SAVVY-Bench. Studies with ground-truth floormaps demonstrate the substantial potential of FloorSAV with accurate spatial information.
CommentsProject page: https://byulharang.github.io/FloorSAV/