跨批次边界的信息路由:Lipschitz 多臂老虎机中的内存-批次权衡
Memory--Batch Tradeoffs in Lipschitz Bandits
浏览论文内容
中文总结 AI 辅助
本文研究随机 Lipschitz 多臂老虎机中内存与批次的权衡,提出含新惩罚项的极小极大遗憾刻画,揭示状态宽度与更新深度不可互换,匹配策略可实现相关性能。
中文摘要 AI 辅助
自适应学习既需要保留观测所蕴含信息的状态,也需要基于该状态采取行动的机会。我们研究随机 Lipschitz 多臂老虎机中的这种宽度-深度权衡。每次拉动后,学习者最多保留 W 比特与奖励相关的活跃状态,并将其拉动组织为最多 B 个已提交批次。当 W ≳_d log(eT) 时,我们刻画了对数因子内的极小极大期望伪遗憾,且该下界对所有 W 均成立。除经典的顺序和无限制内存批次成本外,该前沿还包含新的惩罚项 T^((d+2)/(d+3)) (1+(B-1)W)^(-1/(d(d+3))),证明状态宽度与更新深度不可互换。这种交互是一种信息路由约束:在区域尺度 s 下,低遗憾要求已提交的动作记录编码 Θ_d(s^(-d)) 个区域决策,而收集的边界状态最多携带 (B-1)W 比特的熵。匹配策略在保留安全活跃集掩码的同时,流式传输并擦除验证统计量,可在内存中或逐片段执行。该定理在完全顺序特例中恢复了全维最坏情况仅批次前沿和对数内存可达性;静态批次边界与可预测自适应边界匹配。
英文摘要
Lipschitz bandits admit near-optimal regret $\widetilde O_d(T^{(d+1)/(d+2)})$ with little memory under full adaptivity, or with few batches under unrestricted memory. We characterize the minimax expected pseudo-regret over $T$ rounds with $W$ bits of memory and at most $B$ batches in dimension $d$. For every memory budget $W$, we prove a lower bound of order $T^{\frac{d+2}{d+3}} (1+(B-1)W)^{-\frac{1}{d(d+3)}}$. When $W\gtrsim_d\log(eT)$, algorithms with fixed batch boundaries attain, up to logarithmic factors, the larger of this bound and the optimal $B$-batch regret with unrestricted memory. The analysis separates fine-scale comparisons from the regional information needed to allocate their samples at low regret. Attaining $\widetilde O_d(T^{(d+1)/(d+2)})$ regret requires both $B=Ω_d(\log\log T)$ and $(B-1)W=\widetildeΩ_d(T^{d/(d+2)})$. With one bit of memory, minimax regret is $\widetildeΘ_d(TB^{-1/(d+2)})$ under adaptive batch boundaries while $Θ_d(T)$ under fixed boundaries.
发表机构
- Fudan University(复旦大学)
- Shanghai Innovation Institute(上海创新研究院)
机构由 AI 辅助整理,请以论文原文为准。