arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DINOde:用于开放词汇语义分割的连续视觉-文本对齐

DINOde: Continuous Vision-Text Alignment for Open-Vocabulary Semantic Segmentation

Sung-Hoon Yoon, Hoyong Kwon, Changgyoon Oh, Kuk-Jin Yoon

arXiv 2607.21371首次发表:更新:

发表机构

Multimodal Intelligence and Perception Lab., DGIST; Visual Intelligence Lab., KAIST(多模态智能与感知实验室,大邱庆北科学技术院; 视觉智能实验室,韩国科学技术院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对开放词汇语义分割中视觉与文本对齐问题,提出基于ODE的DINOde框架,通过语义文本流和全局上下文流、速度切向投影等组件,连续对齐视觉与文本,实验证明该方法优于现有方法,性能达领先水平。

AI 中文摘要

开放词汇语义分割(OVSS)利用文本语义对预定义类别之外的对象进行分割。虽然自监督模型DINOv3提供了强大的结构化视觉表示,但其缺乏原生文本对齐阻碍了其在OVSS中的直接应用。为弥合这一差距,我们提出了DINOde,一个基于常微分方程(ODE)的框架,它将CLIP文本嵌入与DINO视觉流形连续对齐。我们的方法采用两个互补组件:(i)语义文本流(STF),它通过连续的ODE轨迹将文本嵌入向DINO流形演化;(ii)全局上下文流(GCF),它逐步细化由DINO的CLS令牌携带的整体图像表示。为在演化过程中保留特征空间的超球面几何,我们进一步引入速度切向投影,将学习到的速度场约束到切空间。通过将对齐建模为连续轨迹,DINOde避免了离散MLP投影中固有的流形纠缠,并产生更稳健的跨模态对齐。广泛实验表明,DINOde始终优于现有方法,并在多个OVSS基准测试中取得了领先性能。代码可在指定网址获取。

英文摘要

Open-vocabulary semantic segmentation (OVSS) leverages textual semantics to segment objects beyond predefined categories. While the self-supervised model DINOv3 provides strong structured visual representations, its lack of native textual alignment hinders its direct application to OVSS. To bridge this gap, we propose DINOde, an ODE-based framework that continuously aligns CLIP text embeddings with the DINO visual manifold. Our approach employs two complementary components: (i) Semantic Text Flow (STF), which evolves text embeddings toward the DINO manifold through a continuous ODE trajectory, and (ii) Global Context Flow (GCF), which progressively refines the holistic image representation carried by DINO's CLS token. To preserve the hyperspherical geometry of the feature space during this evolution, we further introduce Velocity Tangent Projection, which constrains the learned velocity field to the tangent space. By modeling alignment as a continuous trajectory, DINOde avoids the manifold entanglement inherent in discrete MLP projections and yields more robust cross-modal alignment. Extensive experiments demonstrate that DINOde consistently outperforms existing methods and achieves state-of-the-art performance across multiple OVSS benchmarks. The code is available at https://github.com/yoon307/DINOde.

CommentsAccepted to ECCV 2026. 27 pages, 8 figures, and 10 tables. Includes supplementary material

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑