arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CapFrame:基于几何伪标签的3D高斯场景中文本引导视角定位

CapFrame: Text-Instructed Viewpoint Grounding in 3D Gaussian Scenes via Geometric Pseudo Labels

Jirong Li, Satoshi Ikehata, Shuhei Kurita, Ikuro Sato

arXiv 2608.30342首次发表:更新:

发表机构

Institute of Science Tokyo; Denso IT Laboratory, Inc.; National Institute of Informatics(东京科学大学; 电装IT实验室; 信息学研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对3D高斯场景中虚拟相机位姿需人工设置的问题,提出CapFrame框架,通过检索-转换-优化流程实现文本引导的视角定位,实验表明其对齐度优于基线方法。

AI 中文摘要

3D高斯溅射(3DGS)技术支持照片级真实感的实时新视角合成,但放置虚拟相机以获取所需帧的过程仍主要依赖人工操作。现有3D场景中基于语言的方法主要聚焦于以物体为中心的定位,即确定要观察的对象,却很少控制其在单帧中的呈现方式,例如主体朝向或帧布局。为解决这一局限,本文提出一项新任务:文本引导视角定位(TIVG),旨在识别3D高斯场景中的6自由度(6-DoF)相机位姿,使渲染帧与文本指令对齐。为完成该任务,本文提出CapFrame,一种部分可微框架,可将语言转换为用于相机位姿优化的几何伪标签。CapFrame遵循检索-转换-优化流程:通过多模态大语言模型(MLLM)的问答评估流程检索相关视角并排序,将指令转换为朝向和布局伪标签,再通过3DGS中的布局损失和朝向损失进行可微优化以细化相机位姿。在38个真实场景和135条指令上开展的实验表明,CapFrame生成的视角与文本的对齐度优于启发式视角搜索和适配轨迹生成基线,该结果通过多模态大语言模型(VLM)指标、MLLM评估和用户研究得到验证。代码可访问:this https URL

英文摘要

3D Gaussian Splatting (3DGS) enables photorealistic real-time novel view synthesis, yet placing a virtual camera to capture a desired frame remains largely manual. Existing language-guided approaches in 3D scenes mainly focus on object-centric grounding, determining what to observe but rarely controlling how it should appear in a single frame, such as subject orientation or frame layout. To address this limitation, we introduce a new task, Text-Instructed Viewpoint Grounding (TIVG), which aims to identify a 6-DoF camera pose in a 3D Gaussian scene whose rendered frame aligns with a text instruction. To solve this task, we propose CapFrame, a partially differentiable framework that converts language into geometric pseudo labels for camera pose optimization. CapFrame follows a Retrieve-Translate-Refine pipeline: it retrieves relevant views and ranks them through a Question-Evaluation process with MLLMs, translates the instruction into orientation and layout pseudo labels, and refines the camera pose via differentiable optimization with layout and orientation losses in 3DGS. Experiments on 38 real-world scenes with 135 instructions indicate that CapFrame produces viewpoints better aligned with texts than heuristic viewpoint search and adapted trajectory generation baselines, validated by VLM metrics, MLLM judges, and user studies. Code is available at: https://github.com/jirongli/CapFrame

CommentsAccepted at ECCV 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑