arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

一种通过逆强化学习实现的用于失真风险度量的抗噪声引出到优化框架

A Noise-Robust Elicit-to-Optimize Framework for Distortion Riskmetrics via Inverse Reinforcement Learning

Yang Liu, Yuhao Liu, Yunran Wei

arXiv 2607.14373首次发表:更新:

发表机构

School of Science and Engineering, The Chinese University of Hong Kong (Shenzhen); School of Mathematics and Statistics, Carleton University(香港中文大学(深圳)理工学院; 卡尔顿大学数学与统计学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出抗噪声引出到优化框架,集成逆强化学习与强化学习。引出方面用自适应贝叶斯IRL方法,优化方面开发无模型RL算法,通过扩展PPO算法优化风险目标,实证研究证明框架在复杂金融环境中的准确性和有效性。

AI 中文摘要

我们提出了一种抗噪声的引出到优化框架,该框架集成了逆强化学习(IRL)和强化学习(RL),用于在以失真风险度量为特征的广泛风险目标下引出代理的风险偏好并优化策略。在引出方面,我们提出了一种自适应贝叶斯IRL方法,从代理的噪声观测决策中推断其潜在风险目标,明确允许代理采取随机和次优行动。我们确定了一组有限的区分问题的存在,这些问题能在候选类中识别出首选的失真风险度量,并证明了算法在一般设置下的收敛速度为$O(\exp(-cm+O(\sqrt{m\log m})))$,其中$c>0$是常数,$m$表示算法迭代次数。在优化方面,我们开发了一种无模型RL算法,用于在条件失真风险度量下优化策略。通过将目标表示为条件成本分位数函数关于失真函数的积分,该方法统一了失真风险度量目标。我们通过用策略、价值和分位数神经网络扩展近端策略优化(PPO)算法来优化各种风险目标,其中分位数网络估计完整的条件成本分位数函数并实现一般风险目标的数值评估。全面的实证研究证明了该框架在复杂金融环境中的引出准确性和有效性。

英文摘要

We propose a noise-robust elicit-to-optimize framework that integrates inverse reinforcement learning (IRL) and reinforcement learning (RL) for eliciting agents' risk preferences and optimizing policies under a broad class of risk objectives characterized by distortion riskmetrics. On the elicitation side, we propose an adaptive Bayesian IRL method that infers agents' latent risk objectives from their noisy observed decisions, explicitly allowing agents to take stochastic and suboptimal actions. We establish the existence of a finite set of distinguishing questions that identifies the preferred distortion riskmetric within the candidate class and prove that the convergence rate of the algorithm is of order $O(\exp(-cm+O(\sqrt{m\log m})))$ under general settings, where $c>0$ is a constant and $m$ denotes the number of algorithm iterations. On the optimization side, we develop a model-free RL algorithm for optimizing policies under conditional distortion riskmetrics. By representing the objective as an integral of the conditional cost quantile function with respect to the distortion function, the method unifies distortion-riskmetric objectives. We optimize diverse risk objectives by extending the Proximal Policy Optimization (PPO) algorithm with policy, value, and quantile neural networks, where the quantile network estimates the full conditional cost quantile function and enables numerical evaluation of general risk objectives. A comprehensive empirical study demonstrates the framework's elicitation accuracy and effectiveness in complex financial environments.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑