arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

提问一无所获:BFCL多轮中行动决策的评分

Asking Earns Nothing: Scoring the Decision to Act in BFCL Multi-Turn

Yangze Liu, Zhongyi Han

arXiv 2610.04429首次发表:更新:

发表机构

Shandong University(山东大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对BFCL多轮基准中应提问轮次被评分忽略的问题,提出基于成对项的决策评分法,发现评分偏向行动而非正确决策,并发布数据与工具。

AI 中文摘要

缺乏所需信息的智能体应提问而非行动,智能体排行榜的任务定义也如此规定。BFCL多轮基准的四个类别中有两个围绕模型应提问的轮次构建,但其评分器从不查看该轮次:该处的黄金轨迹为空,检查器跳过该轮,脚本化用户无法回答,因此提问一无所获,在该轮次猜测也无代价,而两次提问则导致该项丢失。该基准还包含针对该决策的对照实验。一个应提问项是一个基础项,从某一轮中移除一条信息,因此同一请求在同一轮索引出现两次,一次完整一次不完整:在完整项上模型应做出改变世界的调用,在不完整项上应提问。我们对每对项评分一个决策,即模型是否在该轮尝试了改变世界的调用,直接从存储的轨迹中读取,无需LLM评判;始终行动和始终提问的得分均为50。在提出该决策的223对项中,gpt-5.4在83.4%的完整轮次上尝试调用,在78.0%的不完整轮次上克制,七个模型中最佳决策准确率为80.7%;在相同项上,官方评分将其排第六,而将在此处居中的模型排第一。添加一行告诉gpt-5.4不要提问,促使其在对的两侧都行动,因此其决策准确率无明显变化,而其官方评分在两个应提问类别和基础孪生项上上升13.5至23.5分;相反的行,告诉gemma-4-31B-it先提问,使其决策提高4.5分,但官方评分无增益。评分随推动行动的方向变化,而非随决策本身变化。我们发布这些对项、一个无需API密钥即可在任何BFCL输出目录上运行的轮级评分器,以及31个手动验证的坏项。

英文摘要

An agent that lacks the information it needs should ask rather than act, and the task definitions of agent leaderboards say so. BFCL multi-turn builds two of its four categories around a turn on which the model is supposed to ask, and its scorer never looks at that turn: the gold trajectory there is empty, the checker skips it, and the scripted user cannot answer, so asking earns nothing, guessing costs nothing on that turn, and asking twice loses the item. The benchmark also contains the control experiment for that decision. A should-ask item is a base item with one piece of information removed from one turn, so the same request appears twice at the same turn index, once complete and once not: on the first the model should make the call that changes the world, on the second it should ask. We score one decision per pair, whether the model attempted a world-changing call on that turn, read off the stored trajectories with no LLM judge; acting always and asking always both score 50. On the 223 pairs that pose this decision, gpt-5.4 attempts the call on 83.4% of the complete turns and holds back on 78.0% of the incomplete ones, the best decision accuracy of seven models at 80.7%; on the same items the official score ranks it sixth and puts first a model that lands in the middle here. One added line telling gpt-5.4 not to ask pushes it toward acting on both sides of the pair, so its decision accuracy shows no detectable change, while its official score rises by 13.5 to 23.5 points on the two should-ask categories and on the base twins; the opposite line, telling gemma-4-31B-it to ask first, improves its decision by 4.5 points and gains no official score. The score moves with the push toward action, not with the decision. We release the pairs, a turn-level scorer that runs on any BFCL output directory without an API key, and 31 manually verified bad items.

Comments19 pages, 2 figures. Code and data: https://github.com/YangzeLiu/asking-earns-nothing

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑