AI 中文总结
本研究针对微视频推荐中主流方法的语义动态纠缠与流匹配模型忽略时间结构的问题,提出PrismRec框架,通过谱语义分解和上下文校准偏好匹配,在四个数据集上较SOTA提升最高22.65%且效率更优。
AI 中文摘要
微视频推荐旨在从历史交互和多模态视频内容中推断用户偏好,从而识别用户感兴趣的下一个视频。然而,主流方法将帧序列压缩为单一整体表示,使共同塑造用户偏好的稳定视觉语义与动态演化信息发生纠缠。同时,基于扩散和流匹配的推荐器仅将生成过程建立在粗略的行为上下文之上,未将其内部时间结构纳入偏好形成过程。因此,我们提出PrismRec,一种结合谱分解的微视频推荐偏好流匹配框架。类似于棱镜将白光色散为其组成光谱,PrismRec设计了谱语义分解(SSF),通过时间频率域中先验引导的可学习频率掩码,从帧级表示中推导互补的静态语义和动态因素。随后,它提出上下文校准偏好匹配(CPM),根据每个用户的特定敏感性对这些因素进行加权,并将校准后的上下文作为结构化条件注入,以引导匹配轨迹朝向目标表示,使视频内容成为偏好形成的内在驱动因素,而非辅助的侧信息。在来自两个平台的四个数据集上进行的实验表明,PrismRec相较于SOTA基准方法的性能提升最高达22.65%,且在对比方法中具有最低的推理成本和峰值内存。
英文摘要
Micro-video recommendation aims to infer user preferences from historical interactions and multimodal video content, thereby identifying the next video of interest. However, prevailing methods compress frame sequences into a single holistic representation, entangling the stable visual semantics and the evolving dynamics that jointly shape user preferences. Meanwhile, diffusion- and flow matching-based recommenders condition their generation process solely on coarse behavioral context, leaving its internal temporal structure outside preference formation. We therefore propose PrismRec, a Preference Flow Matching framework with Spectral Factorization for Micro-video Recommendation. Analogous to a prism that disperses white light into its constituent spectrum, PrismRec devises Spectral Semantic Factorization (SSF) to derive complementary static semantic and dynamic factors from frame-level representations via a prior-guided learnable frequency mask in the temporal frequency domain. Then, it proposes Context-Calibrated Preference Matching (CPM) to weigh them with each user's specific sensitivity and inject the calibrated context as a structured condition to steer the matching trajectory toward the target representation, making video content as an intrinsic driver of preference formation rather than auxiliary side information. Experiments on four datasets from two platforms show that PrismRec surpasses the SOTA baseline by up to 22.65%, with the lowest inference cost and peak memory among the compared methods.