arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.02152cs.IRcs.AI

超越模态和谐:用于冲突感知多模态推荐的正交纯化与拓扑引导型混合专家模型

Beyond Modality Harmony: Orthogonal Purification and Topology-Guided MoE for Conflict-Aware Multimodal Recommendation

Jialin Liu, Zhaorui Zhang, Ray C. C. Cheung

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对多模态推荐系统的模态-拓扑冲突问题,提出OrthoRec模型,通过CGOP纯化模态特征、TAR-MoE优化门控及safe-SSL目标,在Amazon数据集上提升了推荐性能与鲁棒性。

中文摘要 AI 辅助

多模态推荐系统(Multimodal Recommender Systems, MRSs)通常依赖于有缺陷的“模态和谐”假设,即认为多模态特征本质上是有益的,且与用户的协同交互模式严格对齐。然而,由于存在具有欺骗性的视觉诱饵点击和不匹配的语义,模态-拓扑冲突在现实场景中普遍存在。盲目整合这些带有噪声的模态特征,不可避免地会污染纯净的协同空间,导致严重的表示失真。为解决这一问题,我们提出用于冲突感知多模态推荐的正交纯化与拓扑引导型混合专家模型(Orthogonal Purification and Topology-Guided MoE for Conflict-Aware Multimodal Recommendation, OrthoRec)。OrthoRec的核心是引入协同引导型正交纯化(Collaborative-Guided Orthogonal Purification, CGOP),该方法从几何上将多模态特征解耦为与纯净协同锚点平行和正交的方向;通过采用保留能量的归一化方法自适应截断正交噪声,CGOP在保留模态固有表示能力的同时,修正具有欺骗性的语义方向。此外,我们设计了拓扑感知路由混合专家模型(Topology-Aware Routing Mixture-of-Experts, TAR-MoE);在协同拓扑的引导下,TAR-MoE采用解耦的sigmoid门控机制,打破传统softmax注意力的零和瓶颈,自主确定每个纯化模态的注入规模。最后,我们引入安全对比学习(safe-SSL)目标,动态惩罚矛盾对的强制对比对齐。在三个真实世界的Amazon数据集上进行的实验表明,OrthoRec始终优于具有竞争力的近期基线方法,且在模态噪声和物品稀疏性下表现出更高的鲁棒性。

英文摘要

Multimodal Recommender Systems (MRSs) typically rely on a flawed "modality harmony" assumption, presuming that multimodal features are inherently beneficial and strictly aligned with users' collaborative interaction patterns. However, modality-topology conflicts are ubiquitous in real-world scenarios due to deceptive visual clickbaits and mismatched semantics. Blindly integrating these noisy modalities inevitably pollutes the pristine collaborative space, causing severe representation distortion. To address this, we propose Orthogonal purification and topology-guided MoE for conflict-aware multimodal Recommendation (OrthoRec). At its core, OrthoRec introduces Collaborative-Guided Orthogonal Purification (CGOP), which geometrically decouples multimodal features into directions parallel and orthogonal to a pure collaborative anchor. By adaptively truncating the orthogonal noise with an energy-preserving normalization, CGOP rectifies deceptive semantic directions while preserving the modality's intrinsic representation capacity. Furthermore, we design a Topology-Aware Routing Mixture-of-Experts (TAR-MoE). Guided by the collaborative topology, TAR-MoE employs decoupled sigmoid gating to break the zero-sum bottleneck of traditional softmax attention, autonomously determining the injection scale for each purified modality. Finally, a safe-SSL objective is introduced to dynamically penalize the forced contrastive alignment of contradictory pairs. Experiments on three real-world Amazon datasets show that OrthoRec consistently outperforms competitive recent baselines and exhibits improved robustness under modality noise and item sparsity.

发表机构

  • City University of Hong Kong(香港城市大学)
  • The Hong Kong Polytechnic University(香港理工大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑