发表机构
Peking University(北京大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对现有3D大语言模型压缩形状和微调覆盖语言能力的问题,OctLLM采用显式八叉树占用令牌序列和稀疏八叉树,通过独立可训练分支添加3D容量,在统一多模态模型中取得图像到3D FID降低17.4%等新最先进结果。
AI 中文摘要
现有的3D大语言模型(LLMs)在两个层面上做出妥协:它们将形状压缩为潜在码本索引或坐标文本,这从模型观察中移除了空间结构;并且它们通过微调主干网络来获取3D模态,这会覆盖其通用语言能力。我们提出OctLLM,解决了这两个局限性。几何信息以八叉树占用令牌的显式3D序列形式进入。然而,完整的八叉树序列随深度迅速增长;因此,OctLLM随机清空倒数第二层节点并省略后代节点,同时保留形状,从而产生一个更短的、基于坐标和深度的稀疏八叉树(S-Octree),用于位置感知的掩码建模生成和3D理解。另一方面,现有方法通过全量微调或LoRA引入新模态,但全量微调成本高昂,LoRA限制了3D容量,且两者都修改了语言通路。OctLLM相反,在与预训练参数分离的参数中添加3D容量:网格令牌通过一部分块中的独立可训练分支路由,而文本和图像令牌保留冻结的视觉-语言通路,两个流通过共享的自注意力进行交互。它训练的参数远少于全量微调,但在统一多模态LLM中创下了新的最先进水平,将图像到3D的FID降低了17.4%,并在渲染接地字幕上比ShapeLLM-Omni提高了28.7个百分点,同时在通用语言基准上与主干网络相匹配。
英文摘要
Existing 3D large language models (LLMs) compromise on two fronts: they compress shapes into latent codebook indices or coordinate text, which removes spatial structure from what the model observes, and they acquire the 3D modality by fine-tuning the backbone, which overwrites its general language ability. We present OctLLM, which addresses both limitations. Geometry enters as an explicit 3D sequence of octree occupancy tokens. However, full octree sequences grow rapidly with depth; OctLLM therefore randomly empties penultimate-level nodes and omits descendants while preserving shape, yielding a shorter coordinate- and depth-anchored Sparse Octree (S-Octree) for position-aware mask-modeling generation and 3D understanding. On the other front, existing methods introduce a new modality with full fine-tuning or LoRA, but full fine-tuning is costly, LoRA limits 3D capacity, and both modify the language pathway. OctLLM instead adds 3D capacity in parameters separate from the pretrained ones: mesh tokens are routed through independent trainable branches in a subset of blocks while text and image tokens retain the frozen vision-language pathway, and the two streams interact through shared self-attention. It trains far fewer parameters than full fine-tuning, yet sets a new state of the art among unified multimodal LLMs, lowering image-to-3D FID by $17.4\%$ and raising render-grounded captioning by $28.7$ points over ShapeLLM-Omni, while matching the backbone on general language benchmarks.
CommentsProject Page: https://plurato.github.io/OctLLM-page/ Code: https://github.com/octree-nn/octllm