UserProxyBench:用于智能体基准测试与训练的LLM用户模拟器评估
UserProxyBench: Evaluating LLM User Simulators for Agent Benchmarks and Training
AI总结:
本研究提出UserProxyBench评估层和用户保真度评分,用于衡量智能体基准测试中模拟用户对指令的遵循程度,并发现过早披露是主要失败模式,同时识别出成本-保真度前沿以指导模拟器选择。
AI中文摘要:
交互式智能体基准测试和多轮强化学习越来越多地将第二个语言模型置于用户角色中。该模拟用户控制着智能体接收信息的时机与内容,然而当前基准测试仅对智能体进行评分,并未直接衡量用户是否正确执行了其被分配的角色。我们引入了UserProxyBench,这是tau-bench系列之上的一个评估层,以及用户保真度评分(UFS),该评分使用基于任务的、独立于智能体成功与否的评分标准,来衡量对基准测试私有用户指令的遵循程度。在375个企业任务中,将智能体固定为GPT-5.5并仅更换用户代理,平均任务奖励变化了15.2个点,而24.4%的成功回合包含用户规范违规。最主要的失败是过早披露:用户在信息被请求之前就提供了信息。这种行为对任务奖励影响甚微,但在成功回合中,它导致智能体平均减少1.06次工具调用,从而在保持奖励的同时改变了被评估的交互。最后,在七个代理中,我们识别出一个经验性的成本-保真度前沿,使从业者能够选择满足所需保真度级别的最便宜的模拟器。
英文摘要:
Interactive agent benchmarks and multi-turn reinforcement learning increasingly place a second language model in the role of the user. This simulated user controls what information the agent receives and when, yet current benchmarks score only the agent and do not directly measure whether the user correctly executed its assigned role. We introduce UserProxyBench, an evaluation layer over the tau-bench family, and the User Fidelity Score (UFS), which measures adherence to the benchmark's private user instructions using task-grounded rubric criteria scored independently of agent success. Holding the agent fixed at GPT-5.5 and varying only the user proxy across 375 enterprise tasks changes mean task reward by 15.2 points, while 24.4% of successful episodes contain a user-specification violation. The dominant failure is premature disclosure: users provide information before it is requested. This behavior has little effect on task reward, yet among successful episodes it causes the agent to make 1.06 fewer tool calls on average, changing the interaction being evaluated while preserving the reward. Finally, across seven proxies we identify an empirical cost-fidelity frontier, enabling practitioners to select the least expensive simulator that satisfies a required fidelity level.