发表机构
UMass Amherst(马萨诸塞大学阿默斯特分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出两种基于内部表示的方法(LTD和ARP)来校准多轮智能体置信度,在三个基准和三个模型上优于现有基线,实现零开销可靠性监控。
AI 中文摘要
随着智能体系统在安全关键应用中迅速普及,衡量与智能体动作相关的置信度变得至关重要。与传统机器学习系统相比,智能体工作流具有复杂的失败模式,涉及规划、工具调用和动态环境交互。在本文中,我们研究了在多轮智能体设置中,模型的内部表示是否为最终任务成功提供更强的信号。我们引入了两种互补方法:潜在轨迹动态(LTD),它总结了交互轨迹中残差流表示的变化;以及动作表示探针(ARP),它根据动作决策时形成的表示预测成功。在三个交互基准(Bash、SQL、Python)和三个模型家族(Qwen14B、Qwen7B、DeepSeek6.7B)中,我们的方法持续优于表面层面的生成和基于序列的校准基线,提供了一种零开销的可靠性监控器,既不需要提示修改,也不需要多样本采样。
英文摘要
As agentic systems getting adopted rapidly in safety critical applications, it is vital to measure the confidence associated with the agentic actions. In comparison to the traditional machine learning systems, agentic workflows have complex failure modes with planning, tool invocation and dynamic environment interactions. In this paper, we investigate whether model's internal representations provide stronger signals of eventual task success in multi-turn agentic setups. We introduce two complementary methods: Latent Trajectory Dynamics (LTD), which summarizes changes in residual-stream representations across an an interaction trajectory, and the Action Representation Probe (ARP), which predicts success from representations formed at action decisions. Across three interactive benchmarks (Bash, SQL, Python) and three model families (Qwen14B, Qwen7B, DeepSeek6.7B), our methods consistently outperform surface level generation and sequence-based calibration baselines providing a zero-overhead reliability monitor that requires neither prompt alterations nor multi-sample rollouts.