arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.11384cs.AI

环境反馈建模很重要:重新思考智能体事后自蒸馏中的反馈处理

Environmental Feedback Modeling Matters: Rethinking Feedback Treatment in Agentic Hindsight Self-Distillation

Hangxi Guo, Fengyuan Liu, Yue Wang, Yuhua Qi, Haoyi Xiong, Fei Sun, Mengnan Du

首次发表
浏览论文内容

中文总结 AI 辅助

本文针对强化学习奖励不可用的问题,提出联合优化环境反馈建模与事后自蒸馏的SELF框架,在τ-bench和AppWorld任务上均优于SDPO、GRPO,提升了智能体能力。

中文摘要 AI 辅助

强化学习常用于在交互环境中训练语言智能体,但在奖励不可用时无法直接应用。近期方法将环境反馈作为事后自蒸馏的特权上下文,但分析表明仅将教师模型以反馈为条件是不够的,因此需重新思考环境反馈在智能体自蒸馏中的使用方式。鉴于环境反馈包含对环境如何响应智能体动作的丰富监督信息,本文提出智能体自蒸馏与环境反馈建模框架(SELF),该框架联合优化环境反馈建模与事后自蒸馏。SELF在将反馈条件化的自教师指导蒸馏到策略的同时,学习预测环境响应。分析揭示了一种相互增强机制:环境反馈建模强化事后监督与策略学习,而自蒸馏提升模型对环境反馈的建模能力。使用Qwen3-8B时,SELF在τ-bench成功率上分别优于SDPO和GRPO 6.4和4.1个百分点,在AppWorld任务目标完成度上分别优于10.71和3.57个百分点。这些结果表明,SELF在智能体自蒸馏中更有效地利用环境反馈,提升了智能体能力。

英文摘要

Reinforcement learning is commonly used to train language agents in interactive environments, but cannot be directly applied when rewards are unavailable. Recent methods use environmental feedback as privileged context for hindsight self-distillation, but our analysis suggests that simply conditioning the teacher on feedback is insufficient, motivating us to rethink how environmental feedback is used in agentic self-distillation. Given that environmental feedback contains rich supervision for modeling how the environment responds to agent actions, we introduce \textit{agentic SElf-distilLation with environmental Feedback modeling} (SELF), a framework that jointly optimizes environmental feedback modeling and hindsight self-distillation. SELF learns to predict environmental responses while distilling guidance from a feedback-conditioned self-teacher into the policy. Our analysis reveals a mutually reinforcing mechanism: environmental feedback modeling strengthens hindsight supervision and policy learning, while self-distillation enhances the model's ability to model environmental feedback. With Qwen3-8B, SELF outperforms SDPO and GRPO by 6.4 and 4.1 percentage points in $τ$-bench success rate, and by 10.71 and 3.57 percentage points in AppWorld task goal completion, respectively. These results show that SELF uses environmental feedback more effectively within agentic self-distillation, improving agent capabilities.

发表机构

  • The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
  • Shanghai AI Laboratory(上海人工智能实验室)
  • University of Science and Technology of China(中国科学技术大学)
  • Institute of Computing Technology, CAS(中国科学院计算技术研究所)

机构由 AI 辅助整理,请以论文原文为准。

↑