arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

审核分布强化学习的风险声明

Auditing the Risk Claims of Distributional Reinforcement Learning

Hari Prasad

arXiv 2607.11607首次发表:更新:

发表机构

Hari Prasad(Hari Prasad)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究分布强化学习智能体风险声明真实性,结合决策指标与统计工具审核,发现多数声明被驳斥,风险反映训练假象,阳性对照证实部分真实声明,训练和集成无法消除问题,重新校准仅否定声明,还记录审核陷阱。

AI 中文摘要

分布强化学习智能体学习完整的回报分布,这些分布越来越被直接视为可解释性、风险敏感控制和安全监控的依据。我们提出一个理论预期但未直接测量的问题:训练后的分布智能体的风险声明是否真实?我们的审核结合了与决策相关的筛选指标(前两个动作之间的超额瓦瑟斯坦差距,等于违反一阶随机优势的质量)、快照重启蒙特卡罗的真实情况以及统计工具(排列空值、自举重驳斥、错误发现率控制),否则审核本身会得出错误结论。在MinAtar上对QR-DQN、C51和IQN进行33次运行,95%置信度下40%-95%最强的风险权衡声明被驳斥,最强声明的位置与盲目声明在统计上无法区分,几乎没有声明可被证实。已知大小的阳性对照证实了96%-100%的真实声明。按照最受关注状态下头部的条件风险价值建议行动,效果从有益到比随机情况差得多。训练风险或集成都无法消除这个问题,重新校准只是通过否定声明来通过审核,头部信息无用,不仅仅是校准错误。我们发布了工具包并记录了两个导致我们自己进行令人信服但错误审核的潜在陷阱。

英文摘要

Distributional reinforcement learning agents learn full return distributions that are increasingly read at face value: for interpretability, risk-sensitive control, and safety monitoring. We ask a question theory anticipates but that has not been measured directly: are the risk claims of a trained distributional agent true? Our audit combines a decision-relevant screening metric (the excess Wasserstein gap between the top two actions, which equals the mass by which first-order stochastic dominance is violated), ground truth from snapshot-restart Monte Carlo, and a statistical harness (permutation nulls, bootstrap refutation, FDR control) without which the audit itself manufactures false conclusions. Across QR-DQN, C51, and IQN on MinAtar (33 runs), 40-95% of the strongest claimed risk trade-offs are refuted at 95% confidence, the placement of the strongest claims is statistically indistinguishable from truth-blind, and essentially no claim is confirmable: for these agents, the learned "risk" reflects a training artifact rather than environment stochasticity. The artifact is structural (fully formed early in training, uncorrelated with final score, idiosyncratic to each seed) and appears unchanged at full-Atari scale, with every top Breakout claim of a pretrained near-state-of-the-art QR-DQN refuted. Positive controls of known magnitude confirm 96-100% of real claims (correlation 0.89-0.92): the reading measures the agents, not the audit. Acting on the heads' CVaR advice at their most-flagged states ranges from beneficial to significantly worse than chance. Neither training for risk nor ensembling removes the artifact, and recalibration passes the audit only by nullifying the claims: the head is uninformative, not merely miscalibrated. We release the toolkit and document two silent pitfalls that produced convincing but wrong audits of our own.

Comments25 pages, 8 figures, 3 tables (main text); includes supplementary material

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑