AI 中文总结
研究探索性均衡中反馈循环问题,通过分析其导数受反向沃尔泰拉 - 抛物型预解式控制,利用块分解确定放大源,得出固定温度下局部均衡分支性质及相关转变,数值计算验证了速率。
AI 中文摘要
熵正则化在时间不一致的随机控制中平滑均衡策略。在低温下,相同的吉布斯响应会强烈放大学习到的奖励和动态中的误差。我们表明探索性均衡的导数由反向沃尔泰拉 - 抛物型预解式控制。沿对齐的正模式,下界具有相同的指数阶。块分解确定了放大源:因果路径贡献\(1 / \tau\)的幂,而正反馈循环可产生指数增长。在固定温度下,局部均衡分支关于有限维模型参数是二次可微的,这产生了函数值德尔塔方法。有界一致椭圆扩散在每个有限维中实现了这种路径 - 循环区分。闭合一个正循环会将根\(n\)线性响应边界从幂律变为\(1 / \log n\)阶;沿循环佩龙模式,当\(N\tau^2 \to \infty\)时,右端点离散化相对一致。仿射模型在兰伯特 - \(W\)温度\(\beta T / W(\beta T\sqrt{n})\)处也给出了精确的非线性转变。数值计算说明了这些速率。
英文摘要
Entropy regularization smooths equilibrium policies in time-inconsistent stochastic control. At low temperature, the same Gibbs response can strongly amplify errors in learned rewards and dynamics. We show that the derivative of an exploratory equilibrium is governed by a backward Volterra-parabolic resolvent. Along an aligned positive mode, a lower bound has the same exponential order. A block decomposition identifies the source of the amplification: causal paths contribute powers of 1/tau, whereas a positive feedback cycle can produce exponential growth. At fixed temperature, a local equilibrium branch is twice differentiable with respect to finite-dimensional model parameters, which yields a function-valued delta method. A bounded uniformly elliptic diffusion realizes this path-cycle distinction in every finite dimension. Closing one positive cycle changes the root-n linear-response boundary from a power law to order 1/log n; along the cyclic Perron mode, right-endpoint discretization is relatively consistent exactly when N tau^2 -> infinity. An affine model also gives an exact nonlinear transition at the Lambert-W temperature beta T / W(beta T sqrt(n)). Numerical calculations illustrate these rates.
Comments36 pages, 2 figures. Includes supplementary material