arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.05878cs.CV

MAVISEG:用于扩散Transformer中零样本开放词汇分割的流形传播与视觉原型

MAVISEG: Manifold Propagation and Visual Prototypes for Zero-Shot Open-Vocabulary Segmentation in Diffusion Transformers

Rajatsubhra Chakraborty, Xujun Che, Ritabrata Chakraborty, Xi Niu, Depeng Xu

首次发表
浏览论文内容

中文总结 AI 辅助

MAVISEG是一种无需训练的优化层,通过恢复扩散Transformer的结构化信号,在六个基准测试的无需训练方法中取得了最强整体结果,其mIoU在各基准中均最优。

中文摘要 AI 辅助

文本到图像的扩散Transformer通过学习生成对象和场景来获取相关知识,使其成为无需训练的零样本开放词汇语义分割的有力候选模型。当前最先进的归因方法会对每个像素进行独立评分,将其特征与固定的文本派生类表示进行比较,无论是输出空间相似度还是交叉注意力权重。这种做法丢弃了模型自身呈现的结构化信号:生成轨迹的时间结构、每个概念的视觉外观统计数据,以及图像自身的成对特征几何结构。我们提出了MAVISEG,这是一种无需训练的优化层,可恢复上述信号。由于其算子仅使用逐像素-概念得分场和像素特征空间,MAVISEG与捕获方式无关,不绑定于某一种归因方法。在六个基准测试中,它在无需训练的方法中取得了最强的整体结果,包括每个基准测试中最佳的mIoU。有趣的是,在初始捕获效果最弱的地方,增益最为显著,且各算子的贡献取决于其优化的得分场中的噪声水平。我们的结果表明,扩散Transformer承载的概念级信息多于当前归因方法所恢复的内容,其中大部分信息在生成掩码的过程中丢失,而非模型本身不存在这些信息。

英文摘要

Text-to-image diffusion transformers learn about objects and scenes by learning to generate them, making them strong candidates for training-free zero-shot open-vocabulary semantic segmentation. State-of-the-art attribution methods score each pixel independently, comparing its features against a fixed text-derived class representation, whether as an output-space similarity or as a cross-attention weight. This discards structured signals the model itself exposes: the temporal structure of the generative trajectory, the visual appearance statistics of each concept, and the image's own pairwise feature geometry. We present MAVISEG, a training-free refinement layer that recovers these signals. Because its operators consume only a pixel-by-concept score field and a pixel feature space, MAVISEG is capture-agnostic rather than tied to one attribution method. Across six benchmarks it achieves the strongest overall results among training-free methods, including the best mIoU on every benchmark. Interestingly, gains are largest where the initial capture is weakest, and individual operators contribute depending on the noise in the field they refine. Our results indicate that diffusion transformers carry more concept-level information than current attribution methods recover, and that much of it is lost on the way to the mask rather than absent from the model.

补充信息

↑