arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

同步多视角神经扩散

Synchronous Multi-view Neural Diffusion

Yongquan Shi, Weijun Huang, Yueyang Pi, Wendi Zhao, Yiqing Shi, Shiping Wang

arXiv 2609.39019首次发表:更新:

发表机构

Fujian Normal University; Fuzhou University(福建师范大学; 福州大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对多视角融合中异步范式限制跨视角交互的问题,提出同步多视角神经扩散SynMDiff,通过联合空间扩散流实现视角内与视角间并发自适应融合,并引入能量拓扑采样与Ego-Net架构提升效率,在真实数据集上大幅超越基线。

AI 中文摘要

多视角学习旨在通过利用不同模态或视角之间的互补性和一致性来学习更全面的表示。然而,现有的多视角融合策略将视角内和视角间的融合视为独立的阶段,没有同时考虑视角内的演化和视角间的依赖关系。这种异步融合范式由于存在冲突的视角特定结构归纳偏差,不可避免地限制了跨视角的交互。因此,信息流在中间路径上容易发生扭曲和压缩,将模型限制在受限的解空间内学习。为了解决这一问题,我们提出了同步多视角神经扩散(SynMDiff),它将多视角特征空间概念化为一个由扩散过程驱动的统一动态系统。通过在联合空间中对任意二元特征交互的扩散流进行建模,SynMDiff实现了视角内和视角间信息的并发和自适应融合。虽然这种同步机制的直接实现会带来高昂的计算成本,但我们进一步引入了基于能量的拓扑采样策略和Ego-Net风格的集中式训练架构,确保了学习和推理过程中的效率和可扩展性。由于其概念上的优雅性和计算上的高效性,在真实世界数据集上的评估表明,SynMDiff大幅优于基线方法。

英文摘要

Multi-view learning seeks to learn more comprehensive representations by exploiting the complementarity and consistency across diverse modalities or views. However, existing multi-view fusion strategies treat intra- and inter-view fusion as independent stages, without simultaneously considering the evolution within views and the dependency across views. Such an asynchronous fusion paradigm inevitably constrains cross-view interactions due to conflicting view-specific structural inductive biases. As a result, information flow is prone to distortion and compression along intermediate pathways, confining the model to learn within a restricted solution space. To address this, we propose Synchronous Multi-view Neural Diffusion (SynMDiff), which conceptualizes the multi-view feature space as a unified dynamical system driven by a diffusion process. By modeling the diffusion flow across arbitrary dyadic feature interactions in a joint space, SynMDiff enables the concurrent and adaptive intra- and inter-view information fusion. While a direct implementation of this synchronized mechanism incurs prohibitive computational costs, we further introduce an energy-based topological sampling strategy and an Ego-Net style centralized training architecture, ensuring both efficiency and scalability during learning and inference. Due to its conceptual elegance and computational efficacy, evaluations on real-world datasets demonstrate that SynMDiff outperforms the baselines by a large margin.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑