arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于加性量化表示的 Bandits 算法

Bandits via Additive Quantized Representations

Ami Tavory, Noam Touitou, Tal Sarig, Frank Cheng, Ido Guy

arXiv 2610.02440首次发表:更新:

发表机构

Meta Platforms(Meta 平台)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出残差量化(RQ)作为表示层,使上下文 Bandits 算法在严格有界内存下实现非线性表达,在 13 个数据集上多数优于非 RQ 方法,并以千倍内存缩减匹配强基线。

AI 中文摘要

上下文 Bandits 需要在非线性奖励建模与在线效率之间取得平衡。树集成和神经网络方法能够捕捉非线性,但需要定期重新训练和大型回放缓冲区。线性模型在每次观测时以 O(1) 内存高效更新,但根本上受限于线性奖励结构。我们提出残差量化(RQ)作为表示层来弥合这一差距。离线训练的 RQ 码本将连续上下文映射为跨多个级别的离散质心分配,并通过影子机制动态设置。这使得一系列加性 Bandit 算法能够以严格有界的内存实现非线性表达能力。在 13 个数据集上,RQ 变体在 11 个数据集上优于其非 RQ 对应方法,通常以较大优势胜出,同时匹配加倍重训练的 XGBoost 和神经网络基线,而内存使用最多减少 1000 倍。

英文摘要

Contextual bandits require balancing nonlinear reward modeling with online efficiency. Tree ensembles and neural methods capture nonlinearities but require periodic retraining and large replay buffers. Linear models update efficiently per observation with O(1) memory, but are fundamentally restricted to linear reward structures. We propose Residual Quantization (RQ) as a representation layer to bridge this gap. An offline-trained RQ codebook maps continuous contexts into discrete centroid assignments across multiple levels, set dynamically through a shadow mechanism. This enables a spectrum of additive bandit algorithms that achieve nonlinear expressivity with strictly bounded memory. Across 13 datasets, RQ variants beat their non-RQ counterparts on 11 of 13 datasets, often by wide margins, while matching doubling-retrain XGBoost and neural baselines using up to 1000 times less memory.

Comments40 pages, 16 figures, 12 tables. Accepted at NeurIPS 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑