AI 中文总结
研究高维域在线比例采样问题,设计支持高效更新和采样的数据结构,在不同对手下证明数据结构深度界,应用于在线学习获高效无悔算法及次线性遗憾保证。
AI 中文摘要
我们研究在σ平滑对手下从高维域进行高效在线比例采样的问题,采样分布由定义在分段结构分区序列上的动态演变权重函数诱导。此设置涵盖广泛应用。核心挑战是随着诱导分区随时间变得日益复杂,维持高效数据结构。我们设计了一种数据结构,支持高效更新和比例采样,避免明确维护指数增长的成本。在σ平滑自适应对手下,证明了数据结构深度的紧密O(√σT)界,在随机顺序对手下为O(log T)界。我们将此框架应用于具有分段结构奖励的在线学习,在全信息和带反馈情况下获得高效无悔算法,具有可证明的次线性遗憾保证。
英文摘要
We study the problem of efficient online proportional sampling from a high-dimensional domain under a $σ$-smoothed adversary, where the sampling distribution is induced by a dynamically evolving weight function defined over a sequence of piecewise-structured partitions. This setting captures a broad range of applications, including principal-agent games (e.g., pricing and contract design), and algorithm configuration and parameter tuning. The central challenge is maintaining an efficient data structure as the induced partition grows increasingly complex over time -- naively, the number of subregions can grow as $O(t^d)$ by round $t$ in $d$ dimensions. We design a data structure that supports efficient updates and proportional sampling while avoiding the cost of explicitly maintaining this exponential growth, where the discontinuities are structured from axis-parallel hyperplanes. Under a $σ$-smoothed adaptive adversary, we prove a tight $O(\sqrt{σT})$ bound on the depth of our data structure, and an $O(\log T)$ bound under a random-order adversary -- to our knowledge, the first such results for this class of problems. We apply this framework to online learning with piecewise-structured rewards, obtaining efficient no-regret algorithms under both full-information and bandit feedback, with provable sublinear regret guarantees.