arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

深度套期保值是强化学习吗?

Is Deep Hedging Reinforcement Learning?

Frédéric Godin

arXiv 2607.13353首次发表:更新:

AI 中文总结

探讨Buehler等人(2019年)的深度套期保值算法是否属于强化学习,针对审稿人质疑,从对强化学习的理解角度出发,指出若按领域标准将蒙特卡罗策略梯度方法等纳入,该算法应属强化学习范畴。

AI 中文摘要

Buehler等人(2019年)的深度套期保值框架通过价格路径的蒙特卡罗模拟和随机梯度下降来训练神经网络策略,以最小化应用于终端套期保值误差的风险度量。我和共同作者将此技术称为强化学习(RL),但遭到了一些审稿人的质疑,理由包括:终端日期才有反馈,无中间奖励信号;缺少价值函数、贝尔曼方程、时间差分(TD)学习和明确探索机制。我认为这些反对意见基于对RL过于狭隘、以TD为中心的理解。一旦按照该领域标准参考文献将RL理解为包括蒙特卡罗策略梯度方法和直接(仅策略)策略搜索等,Buehler等人(2019年)的深度套期保值算法就完全属于RL范畴。

英文摘要

The deep hedging framework of Buehler et al. (2019) trains a neural network policy, via Monte Carlo simulation of price paths and stochastic gradient descent, to minimize a risk measure applied to the terminal hedging error. In a recent stream of papers, my coauthors and I have described this technique as reinforcement learning (RL). Several peers have, on occasion, expressed the view that deep hedging does not constitute genuine RL, on two grounds, among others: first, that because feedback is generated only at the terminal date, with no intermediate reward signal, the method cannot constitute genuine RL; and second, that the absence of a value function, a Bellman equation, temporal-difference (TD) learning, and an explicit exploration mechanism disqualifies the method from the RL category altogether, so that it should instead be labeled a neural-network method for stochastic optimal control. The present note argues instead that both objections rest on an unduly narrow, TD-centric reading of what constitutes RL, and that once RL is understood, as it is in the standard references of the field, to include Monte Carlo policy-gradient methods and direct (actor-only) policy search as first-class members, the deep hedging algorithm of Buehler et al. (2019) falls squarely within the RL umbrella.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑