DCM 老虎机:面向多点击的多人信息非对称级联老虎机
DCM Bandits: Multiplayer Information Asymmetric Cascading Bandits for Multiple Clicks
浏览论文内容
中文总结 AI 辅助
本研究将DCM老虎机扩展至多人信息非对称多点击场景,针对含非对称性的三种场景提供次线性悔恨保证,改进了单智能体结果,算法在非对称环境中表现良好,凸显反馈结构的关键作用。
中文摘要 AI 辅助
本研究将依赖点击模型(DCM)老虎机扩展至多人信息非对称场景,多个智能体与共享的排序列表交互,且每个会话可能产生多次点击,这为选择策略带来了新挑战。我们研究了(1)动作和(2)奖励方面的非对称性,针对至少存在一种非对称性的三种场景,提供了次线性悔恨保证。为这些场景建立匹配的信息论下界是一个开放性问题。我们进一步表明,对于小终止概率,无需知晓终止排序,这改进了先前的单智能体结果。实验证实,我们的算法在非对称环境中表现良好,并凸显了反馈结构的关键作用,即全反馈与首次点击反馈的区别,在协调探索和最小化悔恨方面的重要性。
英文摘要
In this work, we extend the Dependent Click Model (DCM) Bandits to a multiplayer information-asymmetric setting, where multiple agents interact with a shared ranked list and may observe multiple clicks per session, introducing new challenges for selection strategies. We study asymmetry in (1) actions and (2) rewards, providing sublinear regret guarantees for three settings where at least one asymmetry is present. Establishing matching information-theoretic lower bounds for these settings is left as an open problem. We further show that for small termination probabilities, the termination ranking need not be known, improving on prior single-agent results. Experiments confirm that our algorithms perform well across asymmetric environments and highlight the critical role of feedback structure, specifically the distinction between full versus first-click feedback, in coordinating exploration and minimizing regret.
发表机构
- University of California, Los Angeles(加利福尼亚大学洛杉矶分校)
机构由 AI 辅助整理,请以论文原文为准。