arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.07475cs.LG

通过有限深度策略敏感性适应智能体行为变化

Adapting to Changes in Agent Behavior via Finite-Depth Policy Sensitivity

Lan Shi, Daigo Shishika, Xuan Wang

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出有限深度策略敏感性框架,通过近似策略Hessian和混合导数估计智能体行为变化的影响,在追逃博弈中实现更优的估计精度与策略适应,并提升零样本回报和微调效果。

中文摘要 AI 辅助

将强化学习策略适应于另一个智能体行为的变化通常需要大量新的交互数据。策略敏感性提供了局部最优策略如何随行为参数变化的一阶预测,但其计算需要二阶导数,且这些导数的影响会跨未来交互传播。我们开发了一个有限深度框架,通过利用参考环境中的信息近似策略Hessian矩阵和混合导数来估计这种敏感性。该方法具有可调节的传播深度,该深度决定导数沿轨迹传播的截断位置。我们刻画了有限深度传播所省略的导数贡献,并推导了近似导数及由此产生的策略敏感性的截断误差界。这些误差界随传播深度增加而非递增,并在全水平传播时消失。以信念驱动的追逃博弈作为验证场景,所提方法通常随着传播深度的增加而实现更低的导数估计误差,并在估计准确性和策略适应性方面均优于基线方法。基于敏感性的初始化相比直接迁移提高了零样本回报,并且在目标环境中的后续微调中也显示出优势。

英文摘要

Adapting a reinforcement learning policy to changes in another agent's behavior typically requires a large amount of new interaction data. Policy sensitivity provides a first-order prediction of how a locally optimal policy changes with a behavioral parameter, but its computation requires second-order derivatives whose effects propagate across future interactions. We develop a finite-depth framework to estimate this sensitivity by approximating the policy Hessian and mixed derivative using information from a reference environment. The method features an adjustable propagation depth which determines where derivative propagation along the trajectory is truncated. We characterize the derivative contributions omitted by finite-depth propagation and derive truncation-error bounds for the approximated derivatives and resulting policy sensitivity. The bounds are nonincreasing with propagation depth and vanish at full-horizon propagation. Using a belief-driven pursuit-evasion game as a validation scenario, the proposed method generally achieves lower derivative-estimation errors as the propagation depth increases and outperforms the baseline methods in both estimation accuracy and policy adaptation. The sensitivity-based initialization improves zero-shot return over direct transfer, and also shows advantages for the subsequent fine-tuning in the target environment.

发表机构

  • George Mason University(乔治梅森大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑