arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Lens3D:面向细粒度3D理解的目标条件视觉注视

Lens3D: Target-Conditioned Visual Foveation for Fine-Grained 3D Understanding

Junming Huang, Zini Chen, Shuaiying Hou, Chi Wang, Qiang Dai, Weiwei Xu

arXiv 2610.06611首次发表:更新:

发表机构

Zhejiang University; LIGHTSPEED(浙江大学; LIGHTSPEED)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

Lens3D通过外部视觉辅助和知识迁移,提升3D大语言模型对细粒度目标的理解,显著改善描述能力并保持原有性能。

AI 中文摘要

现有的3D大语言模型往往忽视细粒度属性和视觉显著性较低的目标及部件,即使场景视频中存在相关证据。我们引入Lens3D,通过外部视觉辅助和知识迁移来提升细粒度目标理解。其LensUnd流程采用3D定位来选择信息丰富、互补的视图,供外部2D视觉语言模型使用,支持细粒度目标描述、小目标定位和细粒度目标问答。LensDistill通过详细描述监督将所得细粒度知识迁移至3D大语言模型,使得无需外部VLM调用即可从原生输入进行描述。我们还构建了LensBench,一个包含2,068个目标、每个目标有三条银标准参考描述的留出评估集。使用Video-3D LLM和3DRS的实验表明,LensDistill显著提升了细粒度目标描述,同时保持了现有的定位和场景级问答性能。这些结果确立了将外部获取的细粒度知识迁移至原生3D大语言模型的可行性。

英文摘要

Existing 3D large language models often overlook fine-grained attributes and less visually salient objects and parts, even when relevant evidence is present in scene videos. We introduce Lens3D to improve fine-grained object understanding through external visual assistance and knowledge transfer. Its LensUnd pipeline adopts 3D localization to select informative, complementary views for an external 2D vision-language model, supporting fine-grained object captioning, small-object grounding, and fine-grained object question answering. LensDistill transfers the resulting fine-grained knowledge to 3D LLMs through detailed caption supervision, enabling captioning from native inputs without external VLM calls. We also construct LensBench, a held-out evaluation set of 2,068 objects with three silver-standard reference descriptions per object. Experiments with Video-3D LLM and 3DRS demonstrate that LensDistill substantially improves fine-grained object captioning while preserving existing grounding and scene-level QA performance. These results establish the feasibility of transferring externally acquired fine-grained knowledge into native 3D LLMs.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑