arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

优化遗憾值

Optimizing Regret

Irene Aldridge

arXiv 2607.18866首次发表:更新:

发表机构

Risk AI Center(风险人工智能中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文基于成本与决策协方差发展协方差遗憾泛函导数理论,推导线性策略梯度,扩展到约束优化等,得出通用最速下降方向,还给出收敛界及算法,应用于投资组合倾斜与大语言模型分配策略。

AI 中文摘要

基于预期遗憾值等于成本与决策之间协方差的恒等式,本文发展了协方差遗憾泛函的完整导数理论。我们推导了加托导数,表明通用最速下降方向是逆势策略$-(c - \bar{c})$,上升则产生动量。对于线性策略$\hat\pi(c) = Ac + b$,梯度是成本协方差矩阵$\Sigma_c$,海森矩阵为零意味着边界最优解,如最小方差投资组合。我们扩展到约束优化、遗憾最小化与阿尔法最大化之间的符号梯度对偶性、与汤普森采样平行的有限样本收敛界,以及仅需输入观测值的梯度下降算法,并应用于投资组合倾斜和基于大语言模型的分配策略。

英文摘要

Building on the identity that expected regret equals the covariance between costs and decisions, this paper develops a derivative theory of the covariance regret functional. We derive the Gâteaux derivative, showing that the universal steepest-descent direction is the contrarian policy $-(c-\bar c)$, while ascent yields momentum. For linear policies $\hatπ(c)=Ac+b$, the gradient is the cost covariance matrix $Σ_c$, with a zero Hessian implying boundary-optimal solutions such as the minimum-variance portfolio. We extend to constrained optimization, sign-gradient duality between regret minimization and alpha maximization, finite-sample convergence bounds paralleling Thompson Sampling, and gradient-descent algorithms requiring only input observations.

Comments12 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑