arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.19914cs.AI

风险下的长期序列决策

Long-Term Sequential Decision Making under Risk

Irmaan, Mirzanejad, Nadjet Bourdache, Abdel-Illah Mouaddib

首次发表
浏览论文内容

中文总结 AI 辅助

研究风险下有限期MDP规划,提出ERQDP方法,无需枚举和采样,通过精确DP解决秩分位数替代问题,能精确评估候选策略,在测试基准中表现良好,可返回认证解或残差差距,支持多种行为。

中文摘要 AI 辅助

我们研究基于根的(坚决的)风险目标下的有限期马尔可夫决策过程(MDP)规划,该目标将秩依赖函数应用于总回报分布。此类目标在回报分布中是非线性的,通常会破坏贝尔曼最优性,因此通过情景树枚举进行直接优化是难以处理的。我们提出了ERQDP,一种无需枚举和采样的方法,通过精确动态规划(DP)解决秩分位数替代问题,通过在离散回报网格上的回报概率质量函数(PMF)上进行DP精确评估候选策略,并在随时循环中细化替代方案,该循环报告目标目标的明确上下差距(证书)直至离散化预算。在测试基准中,ERQDP返回经过认证的解决方案或明确的残差差距,实现快速风险参数扫描并大幅提高运行时收益,并支持风险规避和风险寻求行为。

英文摘要

We study finite-horizon MDP planning under \emph{root-based} (resolute) risk objectives that apply a rank-dependent functional to the distribution of total returns. Such objectives are non-linear in the return distribution and generally break Bellman optimality, so direct optimization by scenario-tree enumeration is intractable. We propose \textbf{ERQDP}, an enumeration-free and sampling-free method that solves a rank--quantile surrogate via exact DP (Dynamic Programming), evaluates candidate policies exactly by DP over return Probability Mass Functions (PMFs) on a discretized return grid (with an explicit rounding bound), and refines the surrogate in an anytime loop that reports an explicit upper--lower gap (certificate) for the target objective up to discretization budgets. Across tested benchmarks, ERQDP returns certified solutions or explicit residual gaps, enables fast risk-parameter sweeps with substantial runtime gains, and supports both risk-averse and risk-seeking behaviors.

发表机构

  • Université de Caen Normandie, ENSICAEN, CNRS, Normandie Univ, GREYC UMR 6072(卡昂诺曼底大学、法国国立民用航空学院、法国国家科学研究中心、诺曼底大学、GREYC实验室(法国国家科学研究中心6072联合研究单位))

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑