AI 中文总结
本文提出了一种算法,将CDF相关目标的regret最小化复杂度从T^{3/4}降低至T^{7/10},并应用于重复双方面贸易中的利润最大化问题。
AI 中文摘要
我们研究了学习CDF相关目标形式为 [g(x)·P_{X~D}(X≤x)] 的 regrets 最小化问题,其中g是已知的Lipschitz函数,D是未知分布。在每轮t中,学习者选择一个点x_t并观察二元反馈I(X_t≤x_t),其中X_t~D。我们设计了一个算法,其regret为~O(T^{7/10}),优于之前最佳已知界~O(T^{3/4}),表明该类目标的维度诅咒可以至少部分被缓解,尽管与Ω(T^{2/3})下界仍存在差距。作为应用,我们的技术为固定价格的重复双方面贸易中的利润最大化提供了相同的~O(T^{7/10}) regrets 界。
英文摘要
We study regret minimization for learning CDF-related objectives of the form \[ g(x)\cdot\mathbb{P}_{X\sim\mathcal{D}}(X\le x), \] over $[0,1]^2$, where $g$ is a known Lipschitz function and $\mathcal{D}$ is an unknown distribution. At each round $t$, the learner selects a point $x_t$ and observes the binary feedback $\mathbb{I}(X_t\le x_t)$, where $X_t\sim\mathcal{D}$. We design an algorithm achieving regret $\widetilde{\mathcal{O}}(T^{7/10})$, improving over the previous best-known bound of $\widetilde{\mathcal{O}}(T^{3/4})$ and showing that the curse of dimensionality can be at least partially lifted for this class of objectives, though a gap remains with the $Ω(T^{2/3})$ lower bound. As an application, our techniques yield the same $\widetilde{\mathcal{O}}(T^{7/10})$ regret bound for profit maximization in repeated bilateral trade with fixed prices.