不可靠的进度条:LLM 智能体能否在执行过程中可靠地报告任务进度?
The Unreliable Progress Bar: Can LLM Agents Reliably Report Task Progress Throughout Execution?
浏览论文内容
中文总结 AI 辅助
本研究系统评估了LLM智能体在任务各阶段报告进度的可靠性,发现可靠性随阶段变化且多数模型中期失准,最新模型末期保守,并提出了覆盖全程的评估协议,警示框架不应仅依赖状态报告控制流程。
中文摘要 AI 辅助
近期的大型语言模型能够发出任务进度信号,智能体框架利用这些信号来决定任务是应继续还是停止,然而,模型是否能在任务的每个阶段可靠地报告其任务进度,以及其报告在何处及如何失败,尚未得到系统性研究。我们在公共基准 τ²-bench 和 StageIF(一个受控测试平台,其中报告检查点被放置在任务生命周期的各个阶段)上评估了这一能力。两种设置都要求在多个任务阶段进行报告。我们发现,报告可靠性取决于任务所达到的阶段,并且我们测试的几乎所有已部署模型在某些阶段可靠,而在其他阶段不可靠。报告失效的位置并非处处相同。大多数已部署模型一旦工作开始进行,准确性就会下降,并在任务完成后恢复。最新一代模型弥补了任务中期的下降,反而在终点线变得保守。我们的研究揭示了任务进度报告中的能力差距,并提供了一种评估协议,该协议覆盖任务执行的整个过程,以评估智能体运行所依赖的这一能力。研究结果表明,智能体框架不应仅凭模型的状态报告来控制任务流程。
英文摘要
Recent large language models can emit task-progress signals that agent frameworks use to decide whether a task should continue or stop, yet whether a model can reliably report its task progress at every stage of a task, and where and how its reports fail, has not been studied systematically. We evaluate this ability on the public benchmark $τ^2$-bench and on StageIF, a controlled testbed in which reporting checkpoints are placed across the task's lifecycle. Both settings require reports at multiple task stages. We find that reporting reliability depends on the stage a task has reached, and that almost every deployed model we test is reliable at some stages and unreliable at others. Where reporting breaks down is not the same everywhere. Most deployed models lose accuracy once work is under way and recover once the task is done. The newest generation closes that mid-task drop and instead grows conservative at the finish line. Our study exposes a capability gap in task-progress reporting and provides an evaluation protocol that spans the whole course of task execution for this ability on which agent operation depends. The findings indicate that agent frameworks should not control task flow on the strength of the model's state reports alone.
发表机构
- Beihang University(北京航空航天大学)
- National University of Singapore(新加坡国立大学)
机构由 AI 辅助整理,请以论文原文为准。