arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.06565cs.CVcs.AIcs.LG

ELSA3D:用于统一3D理解与生成的弹性语义锚定

ELSA3D: Elastic Semantic Anchoring for Unified 3D Understanding and Generation

Tianjiao Yu, Xinzhuo Li, Yifan Shen, Onkar Susladkar, Yuanzhe Liu, Xiaona Zhou, Ismini Lourentzou

首次发表
浏览论文内容

中文总结 AI 辅助

研究统一3D基础模型文本-3D交互隐式问题,提出ELSA3D模型,通过弹性语义锚定联合构建语言和几何推理,用尺度感知八叉树令牌化器和锚定令牌表示几何,轻量级路由器使计算和推理弹性化,性能领先且减少计算量和延迟。

中文摘要 AI 辅助

统一的3D基础模型期望在单个框架内生成3D资产并进行语言推理,但其文本-3D交互大多是隐式的。现有方法将文本和3D令牌连接成扁平序列并依赖自注意力,将粗略结构线索和精细几何细节合并为无差别的表示。我们引入ELSA3D,一个通过弹性语义锚定解决此问题的统一3D模型,沿着匹配的抽象尺度联合构建语言和几何推理。ELSA3D用尺度感知八叉树令牌化器表示几何,引入锚定令牌,即稀疏跨模态单元,选择语义线索,将其路由到最相关的3D尺度,检索特定尺度的几何证据,并将融合信号写回统一表示,保持交互稀疏而精确。一个轻量级的逐块路由器使计算和推理具有弹性,选择哪些文本令牌在何种几何尺度实例化锚定,使跨模态能力集中在最需要对齐的地方。ELSA3D在图像到3D生成、文本到3D生成和3D字幕方面取得了领先性能,优于最强的统一基线,同时相对于同一模型的非弹性版本,FLOP和推理延迟大致减半。

英文摘要

Unified 3D foundation models aspire to generate 3D assets and reason about them in language within a single backbone, but their text-3D interaction remains largely implicit. Existing methods concatenate text and 3D tokens into a flat sequence and rely on self-attention, collapsing coarse structural cues and fine geometric details into one undifferentiated representation. We introduce ELSA3D, a unified 3D model that addresses this with elastic semantic anchoring, structuring language and geometric reasoning jointly along matched abstraction scales. ELSA3D represents geometry with a scale-aware octree tokenizer and introduces Anchor Tokens, sparse cross-modal units that select semantic cues, route them to the most relevant 3D scale, retrieve scale-specific geometric evidence, and write the fused signal back into the unified representation, keeping interaction sparse yet precise. A lightweight per-block router makes both computation and reasoning elastic, choosing which text tokens instantiate anchors at which geometric scale so that cross-modal capacity concentrates where alignment is most needed. ELSA3D achieves state-of-the-art performance across image-to-3D generation, text-to-3D generation, and 3D captioning, outperforming the strongest unified baseline while roughly halving FLOPs and inference latency relative to the non-elastic version of the same model.

发表机构

  • University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑