arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

当语义一致编码遇见视图-标签异质性建模:不完整多视图多标签学习的统一框架

When Semantically Consistent Encoding Meets View-Label Heterogeneity Modeling: A Unified Framework for Incomplete Multi-View Multi-Label Learning

Chengliang Liu, Bo Li, Bob Zhang, Yanghao Zhou, Jie Wen, Wenwu Wang

arXiv 2609.07525首次发表:更新:

发表机构

University of Macau; Beijing Institute of Technology; Harbin Institute of Technology, Shenzhen; University of Surrey(澳门大学; 北京理工大学; 哈尔滨工业大学(深圳); 萨里大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出V2L统一框架,通过扰动感知编码和主动视图-标签相关性建模,解决不完整多视图多标签学习中的语义一致与视图异质性难题,在五个基准上取得领先性能。

AI 中文摘要

不完整多视图多标签学习不仅需要从部分观测的视图中进行稳健的语义聚合,还需要对视图特定证据进行标签感知的利用。现有方法通常侧重于共享表示学习或决策级融合。前者提高了对缺失视图的鲁棒性,但倾向于将标签判别性的视图特定线索压缩到单个潜在表示中。后者保留了单个视图的预测,但通常依赖于固定或全局学习的融合权重,忽略了不同实例的不同标签可能需要不同的视图。为了解决这些局限性,本文提出了V2L,一个用于不完整多视图多标签分类的统一表示-决策框架。在表示方面,V2L通过扰动感知编码机制从不完整视图中构建语义一致变分后验,提供了稳定的共享语义基础。在决策方面,V2L引入了一种主动的视图-标签相关性建模策略,该策略估计实例级和标签级的视图贡献,使每个标签预测能够自适应地选择有用的视图特定证据。从模型架构的角度来看,这两个重要策略通过混合融合架构集成到一个统一框架中,同时满足跨视图语义一致性和表示互补性的要求。在不完整和完整设置下的大量实验表明,V2L在五个基准上取得了领先的性能。代码可在以下网址获取:此HTTPS URL。

英文摘要

Incomplete multi-view multi-label learning requires not only robust semantic aggregation from partially observed views, but also label-aware exploitation of view-specific evidence. Existing approaches usually emphasize either shared representation learning or decision-level fusion. The former improves robustness against missing views, yet tends to compress label-discriminative view-specific cues into a single latent representation. The latter preserves individual view predictions, but often relies on fixed or globally learned fusion weights, ignoring that different labels of different instances may require different views. To address these limitations, this paper presents V2L, a unified representation-decision framework for incomplete multi-view multi-label classification. On the representation side, V2L constructs semantically consistent variational posteriors from incomplete views through a perturbation-aware encoding mechanism, which provides a stable shared semantic basis. On the decision side, V2L introduces an active view-label relevance modeling strategy that estimates instance-wise and label-wise view contributions, allowing each label prediction to adaptively select useful view-specific evidence. From the perspective of model architecture, these two important strategies are integrated into a unified framework through a hybrid fusion architecture, simultaneously meeting the requirements of cross-view semantic consistency and representational complementarity. Extensive experiments under both incomplete and complete settings show that V2L achieves leading performance on five benchmarks. Code is available at: https://github.com/justsmart/V2L.

CommentsAccepted by IEEE TPAMI

DOI:10.1109/TPAMI.2026.3728832

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑