arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向问题引导的主动视觉的无渲染前瞻方法

Rendering-Free Lookahead for Question-Guided Active Vision

Koya Sakamoto, Daichi Azuma, Shuhei Kurita, Naoya Chiba, Yusuke Iwasawa, Yutaka Matsuo, Taiki Miyanishi

arXiv 2610.11039首次发表:更新:

发表机构

The University of Tokyo; National Institute of Informatics; Institute of Science Tokyo; The University of Osaka(东京大学; 情报学研究所; 东京科学大学; 大阪大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对依赖视角的问题回答任务,提出无渲染前瞻(RFL)策略,通过离线训练学习预测未来可答性,在E3VS-Bench测试集上较基线提升平均评判者评分43%。

AI 中文摘要

主动机器人视觉需要控制相机以获取当前视角下隐藏的任务相关信息,例如确定盒子内部的内容可能需要抬起相机并向下查看。对于依赖视角的问题回答,挑战在于选择能暴露回答问题所需视觉证据的相机运动。尽管视觉语言模型(VLM)可解释观测图像,但选择此类运动需要预判未见过视角的有用性。我们将这种有用性量化为可答性,即VLM对某一视角足以回答问题的估计值,并提出无渲染前瞻(RFL),这是一种通过预测未来可答性对候选相机运动进行排序的视角选择策略。RFL将视觉前瞻从部署阶段转移到离线训练阶段:训练时,特权教师在3D高斯溅射(3DGS)场景中渲染候选未来视角,并使用冻结的VLM计算一步和两步可答性目标;通过两阶段蒸馏,学生模型学习从问题、近期视觉观测和候选相机运动中预测这些动作值。部署时,RFL使用这些预测值选择相机运动,无需渲染未来视角。在未见过环境中的377个E3VS-Bench测试回合上,RFL相比使用相同VLM的直接动作基线,将平均评判者评分提高了43%。这些结果支持从特权视觉前瞻中学习相机控制策略,以用于依赖视角的问题回答。

英文摘要

Active robot vision requires controlling the camera to reveal task-relevant information that is hidden from the current viewpoint. For example, determining what is inside a box may require raising the camera and looking down into it. For viewpoint-dependent question answering, the challenge is to select camera motions that expose the visual evidence needed to answer the question. Although vision-language models (VLMs) can interpret observed images, selecting such motions requires anticipating the usefulness of unseen views. We quantify this usefulness as answerability, a VLM's estimate that a view suffices to answer the question, and present Rendering-Free Lookahead (RFL), a viewpoint-selection policy that ranks candidate camera motions by predicted future answerability. RFL transfers visual lookahead from deployment to offline training. At training, a privileged teacher renders candidate future views in 3D Gaussian Splatting (3DGS) scenes and uses a frozen VLM to compute one- and two-step answerability targets. Through two-stage distillation, a student learns to predict these action values from the question, recent visual observations, and a candidate camera motion. At deployment, RFL uses these predicted values to select camera motions without rendering future views. On 377 E3VS-Bench test episodes in unseen environments, RFL improves the mean judge score by 43\% over a direct-action baseline using the same VLM. These results support learning camera-control policies from privileged visual lookahead for viewpoint-dependent question answering.

CommentsProject page: https://k0uya.github.io/rfl-proj/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑