MURAL:基于自适应边学习的多模态不确定性感知推荐
MURAL: Multimodal Uncertainty-aware Recommendation via Adaptive edge Learning
- American University(美利坚大学)
- Ulsan National Institute of Science and Technology(蔚山国立科学技术院)
- Michigan State University(密歇根州立大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
MURAL是基于自适应边学习的多模态推荐框架,通过动态拓扑发现与不确定性感知融合等技术,在TikTok、Amazon等基准上显著超越现有方法,兼具高准确率、可解释性与鲁棒性。
AI中文摘要:
多模态图神经网络已成为推荐系统的标准方法,其通过内容特征扩充稀疏的交互数据。然而当前架构面临两大瓶颈:一是结构刚性,依赖静态预计算的相似性图,无法适配不断演变的偏好;二是语义脆弱性,会不加区分地融合带噪声的模态信号,从而扭曲协作信号。本文提出MURAL(Multimodal Uncertainty-aware Recommendation via Adaptive edge Learning),这是一个统一框架,将多模态推荐从固定结构扩充转向动态拓扑发现。为解决结构刚性问题,自适应边学习器结合可微检索增强策略与近似最近邻搜索,以发现语义自适应且计算可扩展(时间复杂度为O(NlogN))的潜在物品-物品关联。为解决语义脆弱性问题,不确定性感知融合模块对异构模态的随机不确定性进行建模,动态降低不可靠特征的权重,同时优先考虑高置信度信号以抵御跨模态噪声。本文进一步采用对比师生对齐方法,将模态特定表示锚定到稳定的行为信号,确保优化稳定性且无梯度泄漏。在TikTok和Amazon等大规模基准数据集上的实验表明,MURAL显著超越了结构类和生成类的基线方法,在实现更优准确率的同时,可通过特定领域的模态主导性提供可解释性,并在极端数据损坏下具备鲁棒性。
英文摘要:
Multimodal Graph Neural Networks have become standard for recommendation by augmenting sparse interaction data with content features. Yet current architectures face two bottlenecks: structural rigidity, from a reliance on static precomputed similarity graphs that cannot adapt to evolving preferences; and semantic fragility, where noisy modality signals are indiscriminately fused, distorting the collaborative signal. We propose MURAL (Multimodal Uncertainty-aware Recommendation via Adaptive edge Learning), a unified framework that shifts multimodal recommendation from fixed structural augmentation to dynamic topology discovery. To address structural rigidity, an Adaptive Edge Learner combines a differentiable retrieval-augmented strategy with an approximate nearest neighbor search to discover latent item-item correlations that are both semantically adaptive and computationally scalable (O(NlogN)). To address semantic fragility, an Uncertainty-Aware Fusion module models the aleatoric uncertainty of heterogeneous modalities, dynamically down-weighting unreliable features while prioritizing high-confidence signals as a defense against cross-modal noise. We further employ a contrastive teacher-student alignment that anchors modality-specific representations to stable behavioral signals, ensuring optimization stability without gradient leakage. Experiments on large-scale benchmarks including TikTok and Amazon show that MURAL significantly surpasses both structural and generative state-of-the-art baselines, achieving superior accuracy while offering interpretability through domain-specific modality dominance and robustness under extreme data corruption.