arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

学习具有稳定优化的LLM智能体的扰动鲁棒策略

Learning Perturbation Robust Policies for LLM Agents with Stable Optimization

Pengxin Wang, Yuanzhe LI, Yuxin Ren, Huanrui Yang, Jingdi Chen

arXiv 2609.34064首次发表:更新:

发表机构

University of Arizona(亚利桑那大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对LLM智能体策略对扰动敏感的问题,提出SPrPO方法,在RL训练中引入自适应敏感性感知扰动,在ALFWorld和WebShop上验证了其提升鲁棒性并保持稳定优化。

AI 中文摘要

强化学习(RL)已成为长程大语言模型(LLM)智能体的有效后训练范式。然而,我们发现由此产生的策略可能对各种策略扰动敏感,如隐藏状态噪声、剪枝和量化。在这项工作中,我们研究如何在策略优化过程中提高扰动鲁棒性。我们首先引入扰动鲁棒策略的概念,并分析扰动策略更新保持稳定单调改进的条件。基于此分析,我们提出稳定扰动鲁棒策略优化(SPrPO),它在RL训练期间应用自适应和敏感性感知的扰动。我们在ALFWorld和WebShop上评估SPrPO,并在多种扰动类型和规模上进行系统实验,显示出改进的扰动鲁棒性,同时保持稳定的策略优化。

英文摘要

Reinforcement learning (RL) has become an effective post-training paradigm for long-horizon large language model (LLM) agents. However, we find that the resulting policies can be sensitive to various policy perturbations, such as hidden-state noise, pruning, and quantization. In this work, we study how to improve perturbation robustness during policy optimization. We first introduce the notion of a perturbation robust policy and analyze conditions under which perturbed policy updates preserve stable monotonic improvement. Based on this analysis, we introduce Stable Perturbation-Robust Policy Optimization (SPrPO), which applies adaptive and sensitivity-aware perturbations during RL training. We evaluate SPrPO on ALFWorld and WebShop and conduct systematic experiments across multiple perturbation types and scales, showing improved perturbation robustness while maintaining stable policy optimization.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑