arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

多样化多项式逻辑回归上下文博弈

Diversified Multinomial Logit Contextual Bandits

Heesang Ann, Taehyun Hwang, Min-hwan Oh

arXiv 2607.11684首次发表:更新:

发表机构

Seoul National University(首尔国立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对现有上下文多项式逻辑回归博弈忽略品类内多样性、次模/组合博弈缺乏结构化选择概率的问题,提出多样化多项式逻辑回归(DMNL)上下文博弈及白盒UCB算法OFU-DMNL,给出遗憾界,实验证明其有收益且运行高效。

AI 中文摘要

现有的上下文多项式逻辑回归(MNL)博弈模型关注相关性驱动的选择,但忽略了品类内多样性的潜在好处,而次模/组合博弈在奖励中编码了多样性,但缺乏结构化的选择概率。我们通过多样化多项式逻辑回归(DMNL)上下文博弈弥合了这一差距,它用一个一般的次模多样性函数增强了MNL选择概率,从而在单个模型中形式化了相关性-多样性的权衡。纳入多样性使得精确的MNL品类优化变得棘手。我们提出了一种基于白盒UCB的算法OFU-DMNL,它通过最大化乐观边际收益逐项目构建品类,避免了黑盒优化预言机。我们表明OFU-DMNL实现了至少(1 - 1/(e + 1))-近似遗憾界\(\tilde{O}(d\sqrt{T/K})\),其中d是上下文维度,K是最大品类大小,T是时间范围,并且相对于标准次模基线获得了改进的近似因子。实验证明了持续的收益,并且相对于穷举枚举,具有可比的遗憾但运行时间大大降低。总体而言,DMNL博弈为不确定性下的多样性感知品类优化提供了实用基础,而OFU-DMNL提供了一种统计和计算高效的解决方案。

英文摘要

Existing contextual multinomial logit (MNL) bandits model relevance-driven choice but ignore the potential benefits of within-assortment diversity, while submodular/combinatorial bandits encode diversity in rewards but lack structured choice probabilities. We bridge this gap with the $\textit{diversified multinomial logit}$ (DMNL) contextual bandit, which augments MNL choice probabilities with a generally submodular diversity function, thereby formalizing the relevance--diversity trade-off within a single model. Incorporating diversity renders exact MNL assortment optimization intractable. We propose a $\textit{white-box}$ UCB-based algorithm, $\texttt{OFU-DMNL}$, that constructs assortments item-wise by maximizing optimistic marginal gains, avoids black-box optimization oracles. We show that $\texttt{OFU-DMNL}$ achieves at least a $(1-\frac{1}{e+1})$-$\textit{approximate}$ regret bound $\tilde{O}\left(d \sqrt{T/K}\right)$, where $d$ is the context dimension, $K$ the maximum assortment size, and $T$ the horizon, and attains an improved approximation factor over standard submodular baselines. Experiments demonstrate consistent gains and, relative to exhaustive enumeration, comparable regret with substantially lower runtime. Overall, DMNL bandits provide a practical foundation for diversity-aware assortment optimization under uncertainty, and $\texttt{OFU-DMNL}$ offers a statistically and computationally efficient solution.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑