arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越分类器的模式连通性:来自生成式与对比式模型的证据

Mode Connectivity Beyond Classifiers: Evidence from Generative and Contrastive Models

Chengzheyi Yao, Yongzhao Zhang, Yongding Tian

arXiv 2608.30366首次发表:更新:

发表机构

University of Electronic Science and Technology of China; Delft University of Technology(电子科技大学; 代尔夫特理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文将模式连通性研究拓展至生成式与对比式模型领域,针对DDPM与CLIP架构提出连接算法,首次证实独立训练的DDPM与NanoCLIP模式间存在模式连通性,为理解现代模型损失景观提供新视角。

AI 中文摘要

深度神经网络(DNN)的损失景观呈现出高度复杂且非凸的特性。近期研究揭示了模式连通性现象,表明独立训练的网络模式可通过连续的低损失路径相连。然而,现有模式连通性研究主要局限于基于分类器的模型,现代复杂模型中是否存在类似几何特性仍是开放问题。本文将模式连通性研究拓展至生成式与对比式领域,具体为DDPM与NanoCLIP。针对DDPM与CLIP的独特架构,我们提出一种感知架构的连接构建算法。大量实证结果首次证明,我们成功发现了独立训练的DDPM与NanoCLIP模式间的模式连通性。本研究为理解现代生成式与对比式模型损失景观的几何特性提供了新视角。

英文摘要

The loss landscape of Deep Neural Networks (DNNs) exhibits highly complex and non-convex properties. Recent studies have revealed the phenomenon of mode connectivity, demonstrating that independently trained network modes can be connected via a continuous low-loss path. However, existing mode connectivity research is predominantly confined to classifier-based models, leaving it an open question whether similar geometric properties exist in modern complex models. In this paper, we extend the boundaries of mode connectivity to generative and contrastive domains (specifically DDPM and NanoCLIP). Addressing the unique architecture of DDPM and CLIP, we propose an architecture-aware connection building algorithm. Extensive empirical results demonstrate for the first time that we successfully discover mode connectivity between independently trained DDPM and NanoCLIP modes. Our work provides a novel perspective for understanding the geometric properties of the loss landscapes in modern generative and contrastive models.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑