arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MDP中稀疏策略部署的Bellman认证舍入

Bellman-Certified Rounding for Sparse Policy Deployment in MDPs

Zhaojun Peng

arXiv 2610.00325首次发表:更新:

AI 中文总结

针对MDP中稀疏策略部署,提出Bellman认证舍入方法,利用包络和加权曲率积分提供候选特定保证,显著提升认证覆盖率并降低界限与损失比。

AI 中文摘要

连续策略优化可能将更新分散到许多状态,即使部署仅允许少数完整的状态级更改。我们研究了在有限MDP中,当连续行混合被舍入为稀疏二值策略时,能保留多少折扣回报。策略相关的访问耦合了行编辑,而长视界使全局曲率界限变得保守。从$2d+2$次Bellman求解中,我们推导出可复用的包络,支持舍入前的均匀和候选特定保证。每次交换的秩二有理表示进一步允许沿实现轨迹进行加权曲率积分。我们证明了当预算扩展时,线性维度依赖不可避免,且精确全局曲率阈值化是困难的。候选特定界限将结构化套件上的舍入前认证覆盖率从48.2%提高到74.1%。在$\u03b3=0.95$时,局部积分将耦合实例上的中位界限与损失比从402.3降至2.08。

英文摘要

Continuous policy optimization may spread an update across many states, even when deployment permits only a few complete state-level changes. We study how much discounted return can be retained when continuous row mixtures are rounded to sparse binary policies in finite MDPs. Policy-dependent visitation couples the row edits, while long horizons make global curvature bounds conservative. From $2d+2$ Bellman solves, we derive reusable envelopes that support uniform and candidate-specific guarantees before rounding. A rank-two rational representation of each exchange further permits weighted curvature integration along the realized trajectory. We prove that linear dimension dependence is unavoidable when the budget scales, and that exact global curvature thresholding is hard. Candidate-specific bounds raise pre-rounding certification coverage from $48.2\%$ to $74.1\%$ on the structured suite. At $γ=0.95$, local integration lowers the median bound-to-loss ratio from $402.3$ to $2.08$ on coupled instances.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑