发表机构
Imperial College London; JPMorgan Chase & Co.; University of St. Gallen(伦敦帝国学院; 摩根大通公司; 圣加仑大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对模拟控制策略在真实系统中易因潜在因素变化丧失鲁棒性的问题,提出平稳歧义的建模原则,构建对应模拟器训练策略,在对冲任务中验证其能持续保持鲁棒性。
AI 中文摘要
在模拟中优化的控制策略,当模拟器的参数$x$由有限数据估计且生成的参数不确定性未在模拟中体现时,在真实系统中可能表现不佳。整合此类歧义的常用方法是在随机抽取的$x$值下模拟系统的每条轨迹;由于策略无法观测抽取值,初始时必须选择在众多可能参数值下均表现良好的控制。然而,若策略逐步观测系统,往往能逐渐推断出$x$的值,使歧义消失;随着时间推移,策略会专门适配其对$x$的估计值,进而丧失鲁棒性,这在潜在因素预计会发生变化的诸多真实系统中是不可取的。例如在金融市场中,对冲衍生品收益的策略应保持对波动率制度变化的鲁棒性。为诱导这种持续的鲁棒性,我们提出在模拟器中训练策略,其中歧义随系统状态变化但不会随时间系统性衰减;我们将这一需求形式化为平稳歧义:模拟器应在潜在状态上诱导平稳滤波过程。我们展示了如何构建此类模拟器,并在对冲问题上验证,经平稳歧义训练的策略能随时间保持对潜在因素的鲁棒性,在真实市场数据上表现优异。作为建模原则,平稳歧义为诸多模拟器设计决策提供指导:哪些模型构成真实模拟器、其参数应如何随机化、模拟器与策略应如何初始化。尽管实验聚焦于对冲,但平稳歧义也可能适用于其他由潜在结构变化的外生随机过程驱动的序列控制问题。
英文摘要
Control policies optimized in simulation can perform poorly in the real system when the parameters $x$ of the simulator are estimated from limited data but the resulting parameter uncertainty is not represented inside the simulation. A common way to incorporate such ambiguity is to simulate each trajectory of the system under a randomly drawn value for $x$. Since the policy cannot observe the drawn value, it must initially choose controls that perform well across many possible parameter values. However, if the policy progressively observes the system, it can often gradually infer the value of $x$, so that ambiguity vanishes. Over time, the policy then specializes to its estimate of $x$ and loses its robustness. This is undesirable in many real systems, where latent factors are expected to shift. In financial markets, for example, a policy hedging a derivative payoff should remain robust to changes in the volatility regime. To induce such continual robustness, we propose training policies in simulators where ambiguity varies with the system's state but does not systematically decay over time. We formalize this requirement as stationary ambiguity: the simulator should induce a stationary filter process over the latent state. We show how to construct such simulators and demonstrate, on hedging problems, that policies trained under stationary ambiguity preserve robustness to latent factors over time, leading to strong performance on real market data. As a modeling principle, stationary ambiguity informs many simulator design decisions: which models make realistic simulators, how their parameters should be randomized, and how simulator and policy should be initialized. While our experiments focus on hedging, stationary ambiguity may also be useful for other sequential control problems driven by exogenous stochastic processes with shifting latent structure.