arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于序列推荐中大型预训练多模态嵌入模型的流感知侧适应

Stream-aware Side Adaptation for Large Pre-trained Multimodal Embedding Models in Sequential Recommendation

Junchen Fu, Kaiwen Zheng, Ioannis Arapakis, Wenhao Deng, Xin Xin, Joemon M. Jose, Xuri Ge

arXiv 2607.10909首次发表:更新:

AI 中文总结

针对大型预训练多模态嵌入模型用于序列推荐时因域不对准性能不佳的问题,提出Stresa框架,通过流感知隐藏适配器融合和残差流适配器,有效解锁模型潜力,实验证明其性能优于标准侧适配器和基线。

AI 中文摘要

最近,像Qwen3-VL Embedding这样的大型预训练多模态嵌入模型在序列推荐中展现出强大潜力,但其因域不对准直接使用时性能欠佳。高效的侧适应是有吸引力的解决方案。现有侧适配器常随深度退化,本文提出Stresa框架,引入流感知隐藏适配器融合(SHAF)在融合时保留历史侧记忆,残差流适配器(ReSA)跨层产生选择性残差更新。实验表明Stresa在多个骨干嵌入模型上优于标准侧适配器和基线,凸显了适应大型嵌入模型用于序列推荐的前景。

英文摘要

Recently, large pretrained multimodal embedding models such as Qwen3-VL Embedding have shown strong promise for sequential recommendation, as they provide reusable semantic item representations across modalities and domains. However, directly using these embeddings often leads to suboptimal performance because of domain misalignment. Efficient side adaptation is therefore an attractive solution. Although adapting all backbone layers should help, existing side adapters often degrade with depth, prompting layer dropping despite the loss of useful hidden states. This is due to two major challenges: (1) the lack of modeling in selecting fused representations during residual addition, and (2) the insufficient preservation of earlier representations during progressive sigmoid fusion. This paper therefore asks a practical question: How can we design a side adaptation approach that effectively unlocks the potential of large pre-trained multimodal embedding models? To address this question, we propose Stresa, a stream-aware side-adaptation framework for frozen large pre-trained multimodal embedding models in sequential recommendation. Stresa introduces Stream-aware Hidden-Adapter Fusion (SHAF) to preserve historical side memory during fusion and Residual Stream Adapter (ReSA) to produce selective residual updates across layers. Empirically, Stresa consistently outperforms standard side adapters and state-of-the-art baselines on public datasets across multiple backbone embedding models. These results highlight the promise of adapting large embedding models for sequential recommendation. Our code is publicly available at https://github.com/GAIR-Lab/Stresa.

CommentsAccepted by ACM MM2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑