发表机构
Zhejiang University; LIGHTSPEED(浙江大学; LIGHTSPEED)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Lens3D通过外部视觉辅助和知识迁移,提升3D大语言模型对细粒度目标的理解,显著改善描述能力并保持原有性能。
AI 中文摘要
现有的3D大语言模型往往忽视细粒度属性和视觉显著性较低的目标及部件,即使场景视频中存在相关证据。我们引入Lens3D,通过外部视觉辅助和知识迁移来提升细粒度目标理解。其LensUnd流程采用3D定位来选择信息丰富、互补的视图,供外部2D视觉语言模型使用,支持细粒度目标描述、小目标定位和细粒度目标问答。LensDistill通过详细描述监督将所得细粒度知识迁移至3D大语言模型,使得无需外部VLM调用即可从原生输入进行描述。我们还构建了LensBench,一个包含2,068个目标、每个目标有三条银标准参考描述的留出评估集。使用Video-3D LLM和3DRS的实验表明,LensDistill显著提升了细粒度目标描述,同时保持了现有的定位和场景级问答性能。这些结果确立了将外部获取的细粒度知识迁移至原生3D大语言模型的可行性。
英文摘要
Existing 3D large language models often overlook fine-grained attributes and less visually salient objects and parts, even when relevant evidence is present in scene videos. We introduce Lens3D to improve fine-grained object understanding through external visual assistance and knowledge transfer. Its LensUnd pipeline adopts 3D localization to select informative, complementary views for an external 2D vision-language model, supporting fine-grained object captioning, small-object grounding, and fine-grained object question answering. LensDistill transfers the resulting fine-grained knowledge to 3D LLMs through detailed caption supervision, enabling captioning from native inputs without external VLM calls. We also construct LensBench, a held-out evaluation set of 2,068 objects with three silver-standard reference descriptions per object. Experiments with Video-3D LLM and 3DRS demonstrate that LensDistill substantially improves fine-grained object captioning while preserving existing grounding and scene-level QA performance. These results establish the feasibility of transferring externally acquired fine-grained knowledge into native 3D LLMs.