基于对决反馈的在线凸优化
Online Convex Optimization with Dueling Feedback
浏览论文内容
中文总结 AI 辅助
针对未被探索的对抗凸设置下的在线凸优化问题,提出将对决反馈转化为近似梯度的归约方法,得到首批遗憾保证结果,含不同目标函数下的改进速率。
中文摘要 AI 辅助
我们研究带有对决(成对比较)反馈的在线凸优化问题,其中学习者仅能观测到两个查询点之间的二元偏好。对决反馈在离散或随机设置中已被充分研究,但对抗凸设置仍未被探索。我们提出一种简单的归约方法,将对决反馈转化为近似梯度,从而可使用标准一阶方法。我们证明该归约下遗憾保证可迁移,得到该设置下的首批结果,包括O(T^{3/4})的静态、自适应及动态遗憾。在附加结构下,我们对平滑目标函数得到改进的O(T^{2/3})速率,对强凸函数得到O(√(T log T))速率。
英文摘要
Noisy binary comparison between two candidates is a common interface between human and learning systems, especially in modern large language model (LLM) post-training alignment. We study online convex optimization with dueling (pairwise comparison) feedback, where the learner observes only a binary preference between two queried points. We consider adversarial sequences of convex losses and measure regret with the loss at both queried points, under a comparison link with a known nonzero slope at the origin. We propose a simple reduction that converts dueling feedback into approximate gradients, enabling the use of standard first-order methods. We show that regret guarantees transfer under this reduction, yielding $\mathcal O(T^{3/4})$ static and adaptive regret, and $\mathcal O(T^{3/4}\sqrt{1+P_T/D})$ dynamic regret with unknown comparator path length $P_T$. For strongly convex losses, the static and adaptive bounds improve to $\widetilde{\mathcal O}(T^{2/3})$. For smooth losses, we presents unified dueling ellipsoidal FTRL, and proves $\widetilde{\mathcal O}(T^{2/3})$ static regret, which improves to $\widetilde{\mathcal O}(\sqrt T)$ under additional strong convexity.
发表机构
- Purdue University(普渡大学)
- IIT Indore(印多尔印度理工学院)
- Mila - Quebec AI Institute(米拉-魁北克人工智能研究所)
- McGill University(麦吉尔大学)
机构由 AI 辅助整理,请以论文原文为准。