AI 中文总结
研究针对视频平台个性化视频缩略图生成问题,提出两阶段框架,先通过偏好感知检索选取视觉锚点,再经VLM引导的扩散管道生成缩略图,实验和用户研究证明该方法性能优且能提升用户参与度。
AI 中文摘要
视频缩略图是吸引用户在视频平台上点击的关键因素,且越来越多地得到自动化支持。然而,现有缩略图生成方法通常产生通用结果,忽视了个人偏好的多样性。因此引入个性化视频缩略图生成这一新颖任务,旨在创建符合用户特定偏好的缩略图。这在两方面具有挑战性:识别视觉锚点及生成个性化缩略图。为此提出两阶段框架,将偏好感知检索与可控生成紧密结合。实验表明该方法性能优于基线,用户研究也证明其能提高点击偏好,增强用户参与度。
英文摘要
Video thumbnails are a key factor for attracting user clicks on video platforms, and are increasingly supported by automation. However, existing thumbnail generation methods typically produce generic results shared across users, overlooking the diversity of individual preferences. We therefore introduce personalized video thumbnail generation, a novel task that aims to create thumbnails tailored to user-specific preferences. It is challenging in two aspects: (i) identifying visual anchors (i.e., key frames) from each video to guide the generation, which requires a balance between personalization and informativeness that existing highlight detection methods fail to achieve; and (ii) generating personalized thumbnails that are both visually coherent and faithful to the original video. As a response, we propose a two-stage framework that tightly couples preference-aware retrieval with controllable generation. In the first stage, a personalized highlight retriever captures fine-grained user-video interactions and incorporates video semantics through summarization, enabling the selection of diverse visual anchors aligned with both user preferences and video contexts. In the second stage, a VLM-guided diffusion pipeline transforms these anchors into thumbnails by extracting and injecting semantically grounded visual cues, improving personalization while preserving visual coherence and fidelity. Experiments on two public datasets show our method delivers state-of-the-art performance compared with both retrieval-based and generative baselines. A user study further demonstrates improved click preference, highlighting its effectiveness in enhancing user engagement. The code is available at https://github.com/hezy18/PVTG.