有品味的智能体:在长时程任务中衡量与提升品味
The Tasteful Agent: Measuring and Improving Taste in Long-Horizon Tasks
浏览论文内容
中文总结 AI 辅助
针对长时程任务中智能体决策能力(品味)缺乏衡量的问题,构建自动生成的Taste-Bench基准,发现前沿模型准确率低且推理预算无效,并通过蒸馏训练提升学生模型决策与端到端成功率。
中文摘要 AI 辅助
LLM智能体越来越多地处理长时程任务,它们在此过程中做出的决策,例如测试哪个假设或基于哪个实现进行构建,决定了整个运行的结果。做出这些良好决策的能力正成为工程和研究智能体的关键能力。我们将做出良好长时程决策的能力称为智能体的品味。虽然现有基准衡量智能体在长时程任务上的端到端成功率,但没有一个衡量智能体的品味。为解决这一问题,我们构建了Taste-Bench,一个从智能体在工程和研究任务中产生的轨迹自动构建的品味问题基准。每个问题呈现一个决策分叉,即轨迹中多个方向可用且其中一个方向通向更好结果的点,评估模型在不看到分叉后发生情况的情况下在这些方向中进行选择。我们无需人工标注,即可从同一任务的并行尝试和单条轨迹内的绕路中自动挖掘这些分叉。我们在Taste-Bench上评估了前沿模型,发现最佳模型仅正确回答了59.7%的问题。我们进一步发现,决策证据在轨迹中出现较晚的分叉对每个模型都更难,且更大的推理预算并不能提高准确率。最后,我们表明品味是可以训练的。我们将已看到结果的教师的判断蒸馏到学生模型中,该学生在未见任务上做出更好的决策,并在保留的SWE-bench Pro任务上提高了端到端成功率。
英文摘要
LLM agents increasingly work on long-horizon tasks, and the decisions they make along the way, such as which hypothesis to test or which implementation to build on, determine the outcome of the whole run. Making these decisions well is becoming a key capability for both engineering and research agents. We refer to the ability to make good long-horizon decisions as the taste of an agent. While existing benchmarks measure the end-to-end success of agents on long-horizon tasks, none of them measures the taste of an agent. To address this problem, we build Taste-Bench, a benchmark of taste questions constructed automatically from trajectories that agents produced in engineering and research tasks. Each question presents a decision fork, a point in a trajectory where multiple directions are available and one of them leads to a better outcome, and the evaluated model chooses among these directions without seeing what happens after the fork. We mine these forks automatically from parallel attempts at the same task and from detours inside a single trajectory, without needing human annotation. We evaluate frontier models on Taste-Bench and find that the best model answers only 59.7% of the questions correctly. We further find that forks whose deciding evidence appears later in the trajectory are much harder for every model, and that a larger reasoning budget does not improve the accuracy. Finally, we show that taste can be trained. We distill the judgment of a teacher that has seen the outcome into a student model, and the student makes better decisions on unseen tasks and improves end-to-end success on held-out SWE-bench Pro tasks.
发表机构
- City University of Hong Kong(香港城市大学)
- Microsoft(微软)
机构由 AI 辅助整理,请以论文原文为准。