arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于冷启动评论推荐中汤普森采样的大语言模型衍生先验

LLM-Derived Priors for Thompson Sampling in Cold-Start Comment Recommendation

Eugene Lee, Oseong Choi, Byungsoo Kang, Taeyeong Jang

arXiv 2608.03382首次发表:更新:

发表机构

NAVER WEBTOON(NAVER WEBTOON)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出用大语言模型(LLM)从评论文本提取语义信号生成贝叶斯先验,用于冷启动评论推荐的汤普森采样,经在线A/B/C测试验证其在稀疏反馈场景中效果显著。

AI 中文摘要

多臂老虎机算法,尤其是汤普森采样,被广泛应用于在线推荐中。尽管这些算法具备从在线反馈中自适应调整的能力,但当新引入的臂几乎没有或完全没有交互历史时,它们往往会面临冷启动的局限。在本研究的场景中,候选臂是用户生成的文本评论,在获得足够的交互反馈之前,其语义内容能够揭示标题的吸引力。因此,我们使用大语言模型(LLM)从评论文本中提取语义信号,并将其转换为有信息量的贝叶斯先验,以在早期反馈稀疏时为汤普森采样提供预热启动。为了考虑响应模式中聚合的细分段差异,我们按每个性别-年龄细分段分别维护和更新后验。在一项真实世界的在线A/B/C测试中,我们将均匀先验与两种基于LLM的设计进行了对比:用于人口统计亲和力线索的性别先验,以及用于标题特定身份线索的内容先验。结果表明,基于LLM的先验在反馈稀疏的场景中最为有益——在积累少量交互证据后,增益达到最大——且先验设计会导致漏斗层面的不同效应。我们进一步分析了先验-奖励对齐和人口统计异质性,发现性别先验的点击导向对齐最强,且处理效应在不同人口统计细分段间存在显著差异。这些发现表明,LLM衍生的先验可作为基于文本的老虎机推荐的实用预热启动机制,同时也揭示了部署中的权衡。

英文摘要

Multi-armed bandit algorithms, especially Thompson sampling, are widely used in online recommendation. Despite their ability to adapt from online feedback, these methods often suffer from cold-start limitations when newly introduced arms have little or no interaction history. In our setting, the candidate arms are user-generated textual comments, whose semantic content can reveal a title's appeal before sufficient interaction feedback is available. We therefore use large language models (LLMs) to extract semantic signals from comment text and convert them into informative Bayesian priors that warm-start Thompson sampling under sparse early-stage feedback. To account for aggregate segment-level differences in response patterns, we maintain and update posteriors separately for each gender-age segment. In a real-world online A/B/C test, we compare a uniform prior with two LLM-based designs: a Gender Prior for demographic-affinity cues and a Content Prior for title-specific identity cues. The results show that LLM-based priors are most beneficial in sparse-feedback regimes -- with the largest gains emerging once a small amount of interaction evidence has accumulated -- and that prior design leads to distinct funnel-level effects. We further analyze prior-reward alignment and demographic heterogeneity, finding that click-oriented alignment is strongest for the Gender Prior and that treatment effects vary substantially across demographic segments. These findings suggest that LLM-derived priors can serve as a practical warm-start mechanism for text-rich bandit recommendation, while also revealing deployment trade-offs.

Comments10 pages, 4 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑