arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.01834cs.AIcs.LG

代码掌控模拟,Jev 掌控评估

Code Owns the Simulation, Jev Owns the Evaluation

Yaodong Yang, Hongyao Tang, Yi Ma, Xingyu Fan, Weixun Wang, Jinpeng Li, Tianpei Yang

首次发表
浏览论文内容

中文总结 AI 辅助

本文测试判断模型Jev,发现其在纯评估任务(如认知反思测试)上表现优异,但在需要模拟预测(如对手动作或子目标顺序)时失败;通过代码提供模拟信息,Jev可成为专家控制器。

中文摘要 AI 辅助

诸如 Jev 这样的判断模型,在单次调用中且不附带推理文本的情况下,为每个描述的选项返回一个概率。这使得它们作为智能体的动作选择层颇具吸引力,但尚不清楚哪些决策可以放心地交由它们处理。我们在反思测试、一次性矩阵博弈、文本游戏 ALFWorld 以及机器人控制上对 Jev 进行了测试,并发现了一个清晰的界限。当正确的选项可以根据输入所描述的内容来判断时(我们称之为“评估”),Jev 能够成功。具体而言,它解决了 99% 的反直觉认知反思测试问题。然而,当正确的选项依赖于“模拟”(即预测输入中未包含的内容,例如对手的动作或必须首先完成的子目标)时,它就会失败。在博弈中,Jev 的表现如同其理性对手随机行动一般,表现欠佳,因为对手的动作并未给出。在 ALFWorld 中,Jev 倾向于选择那些提及任务描述中指定对象或地点的命令。例如,给定任务“将一把干净的刀放入抽屉”,Jev 会直接将未清洗的刀带到抽屉,而不是先在洗手池清洗它。令人惊讶的是,许多此类失败并非源于知识缺乏。当被单独询问对手将采取什么行动时,Jev 通常能正确回答,并且在给定对手动作的情况下也能做出良好响应。当一次调用必须同时执行模拟并基于模拟进行评估时,它就会失败。这提示我们让代码来执行预测或模拟。当代码提供这些信息时(例如 ALFWorld 中的前瞻以及机器人控制中的物理模拟),Jev 凭借其通用的评估能力成为专家级控制器。

英文摘要

Judgment models such as \jev{} return, in a single call and without reasoning text, a probability for each described option. This makes them attractive as an agent's action-selection layer, but it is unclear which decisions they can be trusted with. We test \jev{} on reflection tests, one-shot matrix games, the text game ALFWorld and robot control, and find a sharp boundary. \jev{} succeeds when the right option can be judged from what the input describes, which we call \emph{evaluation}. Specifically, it solves 99\% of the counterintuitive Cognitive Reflection Test questions. However, it fails when the right option depends on \emph{simulation} (i.e., predicting something not in the input), such as the opponent's action or the subgoal that must come first. In games, \jev{} plays suboptimally as if its rational opponent acted at random, because the opponent's action is not given. In ALFWorld, \jev{} favors commands that mention an object or place named in the task description. For example, given the task ``put a clean knife in the drawer'', \jev{} carries an unwashed knife straight to the drawer instead of first washing it at the sink. Surprisingly, many of these failures are not due to a lack of knowledge. Asked separately what the opponent will do, \jev{} usually answers correctly, and it responds well given the opponent's action. It fails when one call must both perform the simulation and evaluate based on it. This suggests letting code make the prediction or simulation. When code supplies it, such as a lookahead in ALFWorld and physics simulation in robot control, \jev{} becomes an expert controller through its general evaluation ability.

发表机构

  • The Chinese University of Hong Kong(香港中文大学)
  • Tianjin University(天津大学)
  • Shanxi University(山西大学)
  • CAIR, Hong Kong Institute of Science & Innovation, Chinese Academy of Sciences(中国科学院香港创新研究院人工智能与机器人创新中心)
  • Nanjing University(南京大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑