arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.09115cs.AIcs.CLcs.LG

从不确定性到行动:学习引导LLM智能体

From Uncertainty to Action: Learning to Steer LLM Agents

发表机构新泽西理工学院 · 北卡罗来纳大学教堂山分校 · 加州大学尔湾分校
另 1 家 · 查看机构详情
  • New Jersey Institute of Technology(新泽西理工学院)
  • University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校)
  • University of California, Irvine(加州大学尔湾分校)
  • City University of Hong Kong(香港城市大学)

机构由 AI 辅助整理,请以论文原文为准。

Hanwen Li, Jinhao Duan, Guanhua Zhu, Junchi Lu, Bo Shen, Chenxi Yuan, Kaidi Xu

首次发表
浏览论文内容

中文总结 AI 辅助

针对LLM智能体引导决策,提出基于逐步结果表学习的VoS监控器,利用不确定性识别失败轨迹并决定引导位置,在12种设置中平均提升7.8个百分点。

中文摘要 AI 辅助

引导一个LLM智能体意味着决定是否纠正它、在哪个步骤纠正以及使用哪种机制。不确定性常被用来决定何时纠正智能体,但它能否指导这些决策仍不清楚。我们使用四种机制中的每一种,在每一步非终止步骤分别引导智能体轨迹,并将每条延续运行至完成。由此产生的逐步结果表(SOT)包含来自三个基准和两个智能体的1,864条轨迹的约82,000条反事实延续。它表明不确定性可以识别失败的轨迹,但没有任何单一信号能可靠地定位引导有帮助的步骤。因此,我们提出了VoS(引导价值),一个轨迹级监控器,支持离线或在线使用,它从SOT学习在每个步骤引导的价值,并据此决定在哪里引导。一个危害预算触发器决定是否引导,限制了VoS干扰的成功轨迹的比例。VoS在所有12种基准、智能体以及离线或在线使用的设置中,相较于未修改的执行平均提高了7.8个百分点,并在11种设置中优于五种现有不确定性触发方法中最强的一种,平均提高了2.9个百分点。消融实验表明,基于测量结果进行训练和严格的危害预算都是必不可少的。

英文摘要

Steering an LLM agent means deciding whether to correct it, at which step, and with which mechanism. Uncertainty is often used to decide when to correct an agent, but whether it can guide these decisions remains unclear. We steer agent trajectories separately at every non-terminal step with each of four mechanisms and run each continuation to completion. The resulting stepwise outcome table (SOT) holds about 82,000 counterfactual continuations of 1,864 trajectories from three benchmarks and two agents. It shows that uncertainty can identify failing trajectories, but that no single signal reliably locates the step at which steering helps. We therefore propose VoS (Value of Steering), a trajectory-level monitor, offline or online, that learns from SOT the value of steering at each step and decides where to steer by it. A harm-budgeted trigger decides whether to steer, limiting the fraction of successful trajectories that VoS disturbs. VoS improves on unmodified execution in all 12 settings of benchmark, agent, and offline or online use, by 7.8 points on average, and outperforms the strongest of five existing uncertainty-triggered methods in 11, by 2.9 points on average. Ablations show that training on measured outcomes and a tight harm budget are both essential.

补充信息

↑