arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.16070cs.CLcs.AI

高效多模态生成式推荐与潜在叙事推理

Efficient Multimodal Generative Recommendation with Latent Narrative Reasoning

Chenxing Wang, Nantao Zheng, Hao Miao, Juyuan Wang, Xinke Jiang, Yuchen Fang, Aolin Li, Haijun Wu

首次发表
浏览论文内容

中文总结 AI 辅助

针对情节性内容续写任务,提出NarraLite框架,通过渐进式频谱压缩和潜在叙事推理实现高效多模态生成式推荐,在提升准确性与连贯性的同时优化效率。

中文摘要 AI 辅助

生成式推荐将物品预测重构为语义标识符生成,然而情节性内容引入了一种根本不同的场景,其中目标由叙事演变而非用户偏好决定。该任务要求模型理解多模态故事情节进展,同时应对由冗余视觉上下文和昂贵的显式推理生成所导致的效率挑战。我们提出了NarraLite,一种高效的多模态生成式推荐框架,该框架联合压缩感知与推理。具体而言,渐进式频谱压缩选择性地将长视觉上下文蒸馏为紧凑的叙事相关证据,在保留过渡关键信息的同时减少冗余视觉计算。潜在叙事推理引入上下文路由的潜在推理标记,并将其上下文表示与未来延续语义对齐,从而无需自回归解码文本理由即可实现隐式叙事推断。我们进一步建立了一个与用户无关的多模态基准,用于短视频剧集在UGC、PGC和OOD场景下的续写。大量实验表明,NarraLite在续写准确性、叙事连贯性和鲁棒性方面持续优于现有方法,同时实现了良好的准确性-效率权衡。

英文摘要

Generative recommendation reformulates item prediction as semantic identifier generation, yet episodic content introduces a fundamentally different setting where the target is determined by narrative evolution rather than user preference. This task requires models to understand multimodal storyline progression while addressing the efficiency challenges caused by redundant visual contexts and costly explicit reasoning generation. We propose \textbf{NarraLite}, an efficient multimodal generative recommendation framework that jointly compresses perception and reasoning. Specifically, Progressive Spectral Compression selectively distills long visual contexts into compact narrative-relevant evidence, preserving transition-critical information while reducing redundant visual computation. Latent Narrative Reasoning introduces context-routed latent reasoning tokens and aligns their contextualized representations with future continuation semantics, enabling implicit narrative inference without autoregressively decoding textual rationales. We further establish a user-agnostic multimodal benchmark for short-form drama continuation across UGC, PGC, and OOD settings. Extensive experiments demonstrate that NarraLite consistently improves continuation accuracy, narrative coherence, and robustness over existing approaches, while achieving a favorable accuracy--efficiency trade-off.

发表机构

  • Weixin Group, Tencent(腾讯微信事业群)
  • The Hong Kong Polytechnic University(香港理工大学)

机构由 AI 辅助整理,请以论文原文为准。

↑