Contextual Bandits for Maximizing Stimulated Word-of-Mouth Rewards
最大化激励口碑奖励的上下文赌博机
机构 * University of California, Berkeley(加州大学伯克利分校)
AI总结 提出上下文多臂赌博机框架,通过学习个体溢出概率并排序连接用户,以最大化激励口碑奖励,实验证明考虑溢出异质性可提升目标定位精度。
Comments Presented at the AAAI 2025 Workshop on Bridging the Gap Between AI Planning and Reinforcement Learning (PRL)