arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SpatialQ:基于视觉的多模态大语言模型(MLLM)理解3D Gaussian Splatting场景质量

SpatialQ: Understanding 3D Gaussian Splatting Scene Quality via Visual-based MLLM

Jingxuan Su, Shenglin Wang, Tiesong Zhao, Ge Li, Wei Gao

arXiv 2607.26595首次发表:更新:

发表机构

Peking University; Peng Cheng Laboratory; Fuzhou University(北京大学; 鹏城实验室; 福州大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对3DGS场景质量评估的局限,本文提出SpatialQ框架,通过VGGT编码器与Qwen-MLLM实现结构感知的多模态质量评估,解决传统IQA仅依赖2D线索的问题。

AI 中文摘要

3D Gaussian Splatting(3DGS)已成为用于新视图合成和3D场景重建的有效表示,对可靠的质量评估需求日益增长。与传统图像质量评估(IQA)不同,3DGS场景的质量不仅取决于渲染视图的感知保真度,还取决于空间结构、跨视图一致性等场景级因素。现有IQA方法受限于对2D感知线索的依赖,而通用多模态大语言模型(MLLM)并非为稳定的质量回归设计,可能产生不可靠的判断。为解决这些局限,本文开发了用于3DGS场景理解的多模态质量评估框架:首先,引入3D感知质量表示学习框架,通过为基于VGGT的编码器添加专用质量头,将多视图图像编码为视图特定特征并聚合以捕捉跨视图一致性,同时通过联合建模深度与点云相关结构信息融入几何线索,实现超越外观驱动特征的结构感知质量表示学习;其次,构建基于接地的多模态推理机制,将原始图像、深度图、点云渲染结果及相机参数共同输入基于Qwen的MLLM。

英文摘要

3D Gaussian Splatting (3DGS) has emerged as an effective representation for novel view synthesis and 3D scene reconstruction, creating an increasing demand for reliable quality assessment. Unlike conventional image quality assessment (IQA), the quality of a 3DGS scene depends not only on the perceptual fidelity of rendered views, but also on scene-level factors such as spatial structure and cross-view consistency. Existing IQA methods are limited by their reliance on 2D perceptual cues, whereas general multimodal large language models (MLLMs) are not designed for stable quality regression and may produce unreliable judgments. To address these limitations, a multimodal quality assessment framework is developed for 3DGS scene understanding. First, a 3D-aware quality representation learning framework is introduced by augmenting a VGGT-based encoder with a dedicated quality head. Multi-view images are encoded into view-specific features and aggregated to capture cross-view consistency, while geometric cues are incorporated through joint modeling of depth and point-cloud-related structural information, enabling the learning of structure-aware quality representations beyond appearance-driven features. Second, a grounded multimodal reasoning mechanism is constructed by jointly feeding original images, depth maps, point cloud renderings, and camera parameters into a Qwen-based MLLM.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑