arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

当历史未能成为经验:语言智能体中的动作校准

When History Fails to Become Experience: Action Calibration in Language Agents

Jingyu Liu, Zhiwen Wang, Yuxin Jing, Huanyu Zhou, Yong Liu

arXiv 2610.02769首次发表:更新:

发表机构

Gaoling School of Artificial Intelligence, Renmin University of China; ByteDance; Beijing Key Laboratory of Research on Large Models and Intelligent Governance; Engineering Research Center of Next-Generation Intelligent Search and Recommendation, MOE(中国人民大学高瓴人工智能学院; 字节跳动; 北京大模型与智能治理研究重点实验室; 教育部新一代智能搜索与推荐工程研究中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对语言智能体未能有效利用交互历史的问题,本文提出动作校准方法,通过显式标注动作结果及学习型校准器,提升任务成功率并减少动作重复。

AI 中文摘要

语言智能体应当借鉴先前的尝试和环境反馈,以改进同一任务中的后续决策。然而,提供额外的交互历史有时会降低任务成功率,这表明智能体并未始终有效地利用这些信息。为探究这一局限,我们考察了智能体如何使用历史。我们发现,历史总体上提升了任务完成度,但即使将过去的动作打乱,这一益处仍有很大部分得以保留。破坏动作与观察之间的对应关系仅导致任务成功率的适度下降。因此,我们假设智能体在决定后续步骤时,并未可靠地将过去的动作与其结果联系起来。为验证这一假设,我们明确地将每次返回的观察标注为前一个动作的结果。这一简单的标注在不引入新环境信息的情况下,提升了任务成功率并减少了下一动作的重复。基于这一发现,我们引入了一个学习型校准器,它显式地重新评估过去的动作,并有选择地记录经验以指导后续决策,从而在仅靠结果标注的基础上进一步提升了任务成功率。

英文摘要

Language agents should draw on prior attempts and environmental feedback to improve subsequent decisions within the same task. However, providing additional interaction history can sometimes reduce task success, suggesting that agents do not consistently use this information effectively. To investigate this limitation, we examine how agents use history. We find that history improves task completion overall, yet much of this benefit persists even when past actions are shuffled. Disrupting the correspondence between actions and observations causes only a modest decline in task success. We therefore hypothesize that agents do not reliably connect past actions with their outcomes when deciding how to proceed. To test this hypothesis, we explicitly label each returned observation as the outcome of the preceding action. This simple annotation improves task success and reduces next-action repetition without introducing new environmental information. Building on this insight, we introduce a learned calibrator that explicitly reassesses past actions and selectively records experience to guide subsequent decisions, improving task success beyond outcome labeling alone.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑